EDBT 2026 Demo / reviewers in the wild / expert
Jianshi Tang
dblp:241/8543
· DBLP profile ↗
12ranked-venue papers
0as first author
12since 2021 · last 2026
0000-0001-8369-0067ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 8 · 8 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 4 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | M3DKV: Monolithic 3D Gain Cell Memory Enabled Efficient KV Cache & ProcessingabstractTransformer-based generative large language models (LLMs) have revolutionized natural language processing, yet their quadratic growth in computational complexity in context length creates severe inference bottlenecks. While LLM keyvalue cache (KV cache) enhances decoding efficiency, prolonged contexts infer frequent KV cache reloads that exacerbate memory bandwidth constraints. To address this hardware challenge, we propose M3DKV-a monolithic three-dimensional (3D) gain cell near-memory computing accelerator featuring back-end-of-line (BEOL) cache layers for in-situ KV matrix buffering and computation and a front-end-of-line (FEOL) base layer for full selfattention operations. Through optimized 3D data organization, inter-layer dataflow management, and intelligent computation scheduling, our design achieves $0.29 \mathrm{~TB} / \mathrm{s} /$ core on-die bandwidth while demonstrating $97.03 \times / 268.01 \times$ speedup over GPU/CPU in the decoding stage and $1.72 \times-262.16 \times$ better area efficiency per parameter versus state-of-the-art accelerators. Jiaqi Yang 0009, Yanbo Su, Yihan Fu, Jianshi Tang, Bonan Yan |
ASP-DAC | 4 |
| 2026 | Dual 3T2R Differential SOT MRAM Array for Energy-Efficient In-Memory Computing in Deep Reinforcement Learning
Yiyuan Xiao, Cancheng Xiao, Bingqian Song, Baiyu Su, Jianshi Tang, Tianxiang Nan |
ISCAS | 8 |
| 2026 | MiniBuf: An On-Chip Buffer Allocation Framework Toward Minimizing Buffer Size and Latency for Memristor-Based CNN AcceleratorabstractMemristor-based Convolutional Neural Network (CNN) accelerators have gained considerable attention due to their low latency and high energy efficiency, making them promising candidates for edge acceleration. Alongside statically stored model weights, dynamically generated intermediate feature maps during inference occupy a significant portion of the on-chip buffer capacity and directly affect the efficiency of hardware pipeline execution. However, there is a lack of theoretical analysis and methods for efficiently allocating on-chip buffer for feature map data. To address this gap, this paper implements three innovative aspects. Firstly, a mathematical model is developed to estimate the minimal buffer size required for pipelined inference of CNNs on memristor-based accelerators, offering accurate and swift evaluations of the buffer size requirement. Secondly, based on this model, the paper establishes mathematical conditions for buffer requirements to maintain a blocking-free pipeline during CNN inference, providing theoretical guidance for on-chip buffer allocation strategies. Thirdly, a simulation-in-loop optimization method is proposed to further reduce latency by efficiently increasing the buffer size of critical layers. To validate our proposed model and method, evaluations were conducted on five representative models: ResNet-18, ResNet-50, YOLO-v5, U-Net, and Faster-RCNN-FPN. The results reveal a remarkably low average estimation error of only 2.6% between the mathematical model and the experimentally measured results, with the maximum error still below 10%. Moreover, our simulation-in-loop optimization strategy achieved significant latency reductions ranging from 5.3% to 57.5% across the five models. Ruihua Yu, Chenhuan Zou, Jiaming Li 0006, Bin Gao 0006, Jianshi Tang, He Qian, Huaqiang Wu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2025 | MTVSM-CIM: A Magnified-TMR VC-SOT-MRAM Computing-in-Memory Macro for Edge AIabstractPursuit of energy efficiency and accuracy in intelligent edge devices drives emerging non-volatile memories (NVM) developing more advanced computing-in-memory (CIM) architectures. The application of high performance magnetoresistive random-access memory (MRAM) in CIM has been limited due to its inherent non-ideal characteristics such as low resistance, low switching ratio, high area overhead and low accuracy. In this work, we introduce an innovative weighted two-transistor-one-MTJ (W-2T1R) VC-SOT-MRAM CIM macro that addresses these challenges with: 1) a novel 2T1R cell with high TMR; 2) an excellent approach to assign multi-bit array weight and the redundancy sub-track for linear promotion; 3) a bitwise input sparsity dataflow and architecture to improve energy and latency. Our proposal, excels at high linear computing, with energy efficiencies of 105.5 TOPS/W, area efficiency of 0.73 TOPS/mm2, throughput of 7.2 TOPS and classification accuracy of 91.46% on CIFAR-10 dataset for configurations comprising 6-bit input, 3-bit weight and 6-bit output on VGG-8 model. Bingqian Song, Cancheng Xiao, Fantao Gao, Mengzhu Li, Ziwei Han, Jianshi Tang, Huaqiang Wu, Tianxiang Nan |
ISCAS | 7 |
| 2024 | HXNOR-PBNN: A Scalable and Parallel Spintronics Synaptic Architecture for Probabilistic Binary Neural NetworksabstractThe combination of computing-in-memory (CIM) architecture and emerging non-volatile memory (NVM) is considered as a promising alternative for low-power electronics. Among different emerging NVMs, the insufficient conductance level of magnetoresistive random-access memory (MRAM) is a serious limitation to introduce its high performance into analogue CIM architecture. In this paper, we propose a probabilistic binary neural networks(PBNN) hardware based on voltage-controlled spin-orbit torque MRAM (VC-SOT-MRAM), which employs 2T2MTJ sub-tracks integrated of the probabilistic switching MTJs to achieve the bit-wise weights sampling and activation. Selective writing is discussed to co-optimized. A pipeline readout circuit with hybrid dynamic reference was suggested to improve the precision and throughput. In addition, a weight mapping method is applied for multi-tracks computing in parallel and kernel duplicating to improve the efficiency. Our proposal, simulated in 28 nm technology, achievean energy efficiency of 312.5 bTOPS/W, throughput of 527.3 bGOPS, accuracies of >90% on MNIST with more acceleration (58×), less resource requirement (32×) and high robust row-parallel computing. Cancheng Xiao, Dingsong Jiang, Jianle Liu, Bingqian Song, Jianshi Tang, Huaqiang Wu, Tianxiang Nan |
ISCAS | 6 |
| 2023 | Architecture-circuit-technology co-optimization for resistive random access memory-based computation-in-memory chips
Yuyi Liu, Bin Gao 0006, Jianshi Tang, Huaqiang Wu, He Qian |
Sci. China Inf. Sci. | 3 |
| 2023 | CLEAR: a full-stack chip-in-loop emulator for analog RRAM based computing-in-memory system
Ruihua Yu, Bin Gao 0006, Yiwen Geng, Yuyi Liu, Qingtian Zhang, Jianshi Tang, Hu He 0001, Ning Deng 0008, He Qian, Huaqiang Wu |
Sci. China Inf. Sci. | 8 |
| 2023 | An Error-Free 64KB ReRAM-Based nvSRAM Integrated to a Microcontroller Unit Supporting Real-Time Program Storage and RestorationabstractNonvolatile SRAM (nvSRAM), which integrates the nonvolatile elements with SRAM using a direct bit-to-bit connection has raised much attention in the past few years, owing to its fast parallel data transfer and fast power-on/off speed. However, few nvSRAM macros have been silicon verified to be enacted through the power-failure event. On the other hand, the capacity of fabricated nvSRAM macro is small (~ Kbit) to date, inhibiting its practical application. This study presents a novel ReRAM-based nvSRAM bitcell with improved reliability and scalability. A 64KB nvSRAM macro was designed and integrated into a 32-bit microcontroller unit (MCU). The chip was fabricated using HfOx-based BEOL ReRAM and a 130nm CMOS technology. To pursue fast storage, a write-without-verify scheme is adopted to program ReRAM, measurement results show that the raw bit error rate between the power outages is < 0.1% for the full macro under such constraint. Cryptography and machine learning applications are successfully performed on the MCU system. For the first time, with the help of correction techniques, we achieved an error-free nvSRAM macro that is reliable enough to store/restore programs and demonstrated a real-time robotic control system empowered by the nvSRAM. The proposed nvSRAM macro has the largest capacity to date. Hanwen Gong, Hu He 0001, Liyang Pan, Bin Gao 0006, Jianshi Tang, Sining Pan, Dabin Wu, He Qian, Huaqiang Wu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2021 | An On-chip Layer-wise Training Method for RRAM based Computing-in-memory ChipsabstractRRAM-based computing-in-memory (CIM) chips have shown great potentials to accelerate deep neural networks on edge devices by reducing data transfer between the memory and the computing unit. However, due to the non-ideal characteristics of RRAM, the accuracy of the neural network on the RRAM chip is usually lower than the software. Here we propose an on-chip layer-wise training (LWT) method to alleviate the adverse effect of RRAM imperfections and improve the accuracy of the chip. Using a locally validated dataset, LWT can reduce the communication between the edge and the cloud, which benefits personalized data privacy. The simulation results on the CIFAR-10 dataset show that the LWT method can improve the accuracy of VGG-16 and ResNet-18 by more than 5% and 10%, respectively, with only 25% operations and 35% buffer compared with the back-propagation method. Moreover, the pipe-LWT method is presented to improve the throughput by three times further. Yiwen Geng, Bin Gao 0006, Qingtian Zhang, Yudeng Lin, Jianshi Tang, Huaqiang Wu, He Qian |
DATE | 9 |
| 2021 | Array-level boosting method with spatial extended allocation to improve the accuracy of memristor based computing-in-memory chips
Bin Gao 0006, Jianshi Tang, He Qian, Huaqiang Wu |
Sci. China Inf. Sci. | 4 |
| 2021 | In-memory Learning with Analog Resistive Switching Memory: A Review and PerspectiveabstractIn this article, we review the existing analog resistive switching memory (RSM) devices and their hardware technologies for in-memory learning, as well as their challenges and prospects. Since the characteristics of the devices are different for in-memory learning and digital memory applications, it is important to have an in-depth understanding across different layers from devices and circuits to architectures and algorithms. First, based on a top-down view from architecture to devices for analog computing, we define the main figures of merit (FoMs) and perform a comprehensive analysis of analog RSM hardware including the basic device characteristics, hardware algorithms, and the corresponding mapping methods for device arrays, as well as the architecture and circuit design considerations for neural networks. Second, we classify the FoMs of analog RSM devices into two levels. Level 1 FoMs are essential for achieving the functionality of a system (e.g., linearity, symmetry, dynamic range, level numbers, fluctuation, variability, and yield). Level 2 FoMs are those that make a functional system more efficient and reliable (e.g., area, operational voltage, energy consumption, speed, endurance, retention, and compatibility with back-end-of-line processing). By constructing a device-to-application simulation framework, we perform an in-depth analysis of how these FoMs influence in-memory learning and give a target list of the device requirements. Lastly, we evaluate the main FoMs of most existing devices with analog characteristics and review optimization methods from programming schemes to materials and device structures. The key challenges and prospects from the device to system level for analog RSM devices are discussed. Bin Gao 0006, Jianshi Tang, Meng-Fan Chang, Xiaobo Sharon Hu, Jan Van der Spiegel, He Qian, Huaqiang Wu |
Proc. IEEE | 3 |
| 2021 | Diagonal Matrix Regression Layer: Training Neural Networks on Resistive Crossbars With Interconnect Resistance EffectabstractResistive crossbars implement parallel vector-matrix multiplication (VMM) in analog fashion, and thus enable fast and energy-efficient neuromorphic systems. However, interconnect resistance and resistive switching devices form a complex resistance network with sneak paths. It could result in severe distortions on the output currents. When implementing neural networks, current distortions also cause significant accuracy loss. This article proposes an accurate and computationally efficient model of VMM in resistive crossbars, called diagonal matrix regression (DMR), and incorporates the model into the topology of neural networks as DMR layer (DMRL). Given an m×n crossbar, two diagonal matrices are calculated directly according to the resistance network in a time complexity of only O(m2+n2). No hyper-parameter needs to be determined manually. Modeling of VMM is implemented in a time complexity of only O(mn). DMRL is developed to replace the weight matrix of neural networks so that the effect of interconnect resistance and the sneak path problem are well handled during ex-situ training. Using this technique, for the task of MNIST and fashion-MNIST classification, the accuracy is dramatically restored. Yan Liao, Bin Gao 0006, Jianshi Tang, Huaqiang Wu, He Qian |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |