Chunmeng Dou

dblp:149/4775 · DBLP profile ↗
← Back
14ranked-venue papers
0as first author
14since 2021 · last 2026
0000-0003-2192-9655ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 11 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021
YearPublicationVenuePosition
2026 A BEOL Ferroelectric FET-based Computing Unit for Digital Computing-in-Memory
Junyu Zhu, Zexue Bian, Weizeng Li, Junzhe Shen, Hanghang Gao, Zhidao Zhou, Zhongze Han, Zhi Li 0062, Hongyang Hu, Chunmeng Dou
ISCAS11
2026 An area/energy-efficient RRAM computing-in-memory macro with fully-charge-domain multi-bit computation
Shengzhe Yan, Zhuoyu Dai, Zeyu Guo 0002, Zhaori Cong, Zhihang Qian, Xiangqu Fu, Chunmeng Dou, Dashan Shang, Jinshan Yue
Sci. China Inf. Sci.10
2026 M3CAM: An MLC RRAM-Based Multi-Bit CAM Design Supporting In-Memory Operation of Multi-State Hamming Distance
abstract
The multi-state Hamming distance (MSHD) is a crucial metric for evaluating the similarity of symbolic sequences in data-intensive search applications, such as genomic analysis. MSHD search is performed by comparing inputs against database entries, which can be efficiently accelerated by content-addressable memories (CAMs), such as multi-level cell (MLC) RRAM CAMs. However, designing MLC RRAM-based MSHD CAM faces critical challenges in area efficiency, primarily due to large CAM cells, complex MSHD computing circuits, and bulky variation-compensating input circuits. To address these issues, we propose three techniques to develop M3CAM, an MLC RRAM-based multi-bit CAM (MCAM) supporting in-memory operation of MSHD. First, we propose a 5T1R MCAM cell with MLC RRAM to support dense symbol matching. Second, we propose an MSHD in-memory operation circuit with only three transistors to support dense MSHD computation. Third, we propose a feedback-driven adaptive input DAC to enable minimal area-overhead compensation for RRAM variation. Compared to the state-of-the-art, the proposed M3CAM reduces MCAM cell area by 74.2%, and reduces MSHD computing circuit transistors by 20%, while expanding the MSHD search range to$16\times $.
Tiankuo Zheng, Chenxin Jiang, Yuhao Shu, Chunmeng Dou, Yajun Ha
IEEE Trans. Circuits Syst. I Regul. Pap.5
2025 An Energy-Efficient High-Utilization Hardware Architecture for Attention Mechanism in Transformer using Balanced Systolic Array and Multi-Row Interleaved Operation Ordering
abstract
Transformer-based neural networks have achieved remarkable performance. Designing energy-efficient and high-speed accelerators for the attention mechanism, which dominates the energy and latency in Transformers, has become increasingly significant. Existing attention accelerators commonly use algorithm-hardware co-design to achieve higher energy efficiency and speed. However, deeply customized algorithms make these accelerators dependent on a particular application. Therefore, optimizing hardware architecture is crucial for achieving general-purpose acceleration. We observe two limitations in the hardware architecture of existing attention accelerators. First, the widely used input stationary, weight stationary, and output stationary systolic arrays (SAs) can’t balance data reuse, register saving, and utilization, which hinders to build more energy-efficient and faster SA-based accelerators. Second, layer-by-layer operation ordering introduces high SRAM access overhead of intermediate results. To address the first limitation, we propose the “Balanced Systolic Array”, which improves energy efficiency by 40% compared to conventional systolic arrays and achieves a utilization rate of 99.5%. To address the second limitation, we propose “Multi-Row Interleaved” operation ordering, which reduces the SRAM energy by 31.7% By integrating two techniques, the proposed attention accelerator achieves a 39% improvement in energy efficiency and a 38% enhancement in throughput×energy efficiency compared to previous works.
Haiyang Zhou, Hongyang Hu, Jinshan Yue, Hanghang Gao, Yuanlu Xie, Xiaoxin Xu, Chunmeng Dou, Ming Liu 0022
DAC7
2025 An energy-efficient FeFET-based computing-in-memory macro using BEOL-integrated HZO ferroelectric capacitors
Weizeng Li, Zhidao Zhou, Linfang Wang, Junyu Zhu, Junzhe Shen, Hongyang Hu, Baihan Wang, Zhi Li 0062, Wang Ye, Zhongze Han, Hanghang Gao, Chunmeng Dou
Sci. China Inf. Sci.12
2025 A monolithic 3D IGZO-RRAM-SRAM-integrated architecture for robust and efficient compute-in-memory enabling equivalent-ideal device metrics
Shengzhe Yan, Zhaori Cong, Zhuoyu Dai, Zeyu Guo 0002, Zhihang Qian, Xufan Li, Chuanke Chen, Nianduan Lu, Chunmeng Dou, Guanhua Yang, Xiaoxin Xu, Di Geng, Jinshan Yue, Ling Li 0013, Ming Liu 0022
Sci. China Inf. Sci.11
2025 An RRAM-Based Computing-in-Memory Macro With Low-Power Readout/Hold Circuits and Activation Differential Strategy for AdderNet
abstract
AdderNet is an innovative neural network (NN) structure that substitutes multiplications with additions in convolutional operations, while computing-in-memory (CIM) is an efficient architecture that tackles the memory bottleneck for von Neumann architectures. Previous work has explored the SRAM-based CIM AdderNet circuits and demonstrates high energy efficiency. However, it still suffers low storage density, repetitive readout, and redundant comparisons. In this brief, an RRAM-based CIM macro is proposed for efficient AdderNet with the following innovations. First, RRAM cells are adopted to replace SRAM for high-density weight storage. A low-power readout and hold circuit is proposed to save redundant read power of weight data held for multiple cycles. Second, an 8-bit comparator with an early-stop strategy is proposed to compare 8-bit activations and weights in one cycle. Third, an activation (ACT) differential strategy is proposed to reduce redundant comparisons. The proposed 28-nm RRAM CIM macro achieves 12.8-TOPS/mm2peak area efficiency and 126-TOPS/W peak energy efficiency, which is$3.0\times $and$1.2\times $compared with the state-of-the-art AdderNet CIM macro.
Zhihang Qian, Shengzhe Yan, Zhuoyu Dai, Zeyu Guo 0002, Zhaori Cong, Yifan He 0003, Chunmeng Dou, Feng Zhang 0014, Jinshan Yue, Yongpan Liu
IEEE Trans. Very Large Scale Integr. Syst.7
2025 A High-Density Energy-Efficient CNM Macro Using Hybrid RRAM and SRAM for Memory-Bound Applications
abstract
The big data era has facilitated various memory-centric algorithms, such as the Transformer decoder, neural network, stochastic computing (SC), and genetic sequence matching, which impose high demands on memory capacity, bandwidth, and access power consumption. The emerging nonvolatile memory devices and compute-near-memory (CNM) architecture offer a promising solution for memory-bound tasks. This work proposes a hybrid resistive random access memory (RRAM) and static random access memory (SRAM) CNM architecture. The main contributions include: 1) proposing an energy-efficient and high-density CNM architecture based on the hybrid integration of RRAM and SRAM arrays; 2) designing low-power CNM circuits using the logic gates and dynamic-logic adder with configurable datapath; and 3) proposing a broadcast mechanism with output-stationary workflow to reduce memory access. The proposed RRAM-SRAM CNM architecture and dataflow tailored for four distinct applications are evaluated at a 28-nm technology, achieving 4.62-TOPS$/$W energy efficiency and 1.20-Mb$/$mm2memory density, which shows$11.35\times $–$25.81\times $and$1.44\times $–$4.92\times $improvement compared to previous works, respectively.
Shengzhe Yan, Xiangqu Fu, Zhihang Qian, Zhi Li 0062, Zeyu Guo 0002, Zhuoyu Dai, Zhaori Cong, Chunmeng Dou, Feng Zhang 0014, Jinshan Yue, Dashan Shang
IEEE Trans. Very Large Scale Integr. Syst.9
2025 An RRAM Digital Computing-in-Memory Macro With Dual-Mode Multiplication and Maximum Value Rounding Adder Tree
abstract
Implementing digital computing-in-memory (DCIM) based on resistive memory (RRAM) faces several critical challenges due to the small signal margin, large device variations, and large energy- and area-overhead induced by the digital adder tree (AT). To address these issues, we propose an RRAM DCIM macro based on the standard foundry one-transistor-one-resistor (1T1R) cell array featuring: 1) dual-mode MAC operation for efficiency- or accuracy-oriented optimization; 2) margin-enhanced digitized unit (MEDU) to amplify the signal ratio; and 3) maximum value rounding AT (MVR-AT) to reduce its power- and area-overhead. A test chip is demonstrated using a 180 nm CMOS process to verify the concept. It achieves a peak energy efficiency (EF) of 63.08 TOPS/W in the efficiency-oriented mode and a minimum error rate of 1.58% in the accuracy-oriented mode. Their combination can meet the requirements of different workloads in AI computing tasks to optimize the overall power consumption with negligible accuracy loss.
Wang Ye, Hanghang Gao, Zhidao Zhou, Linfang Wang, Weizeng Li, Zhi Li 0062, Jinshan Yue, Xiaoxin Xu, Hongyang Hu, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.11
2024 A 2T P-Channel Logic Flash Cell for Reconfigurable Interconnection in Chiplet-Based Computing-In-Memory Accelerators
abstract
In this work, we propose a two-transistor (2T) p-type channel (p-channel) logic-compatible flash cell. Compared to the previous designs, the proposed structure features reduced area-cost and enhanced ability to pass through the logic ‘1’. Due to these advantages, we explore its application as the reconfigurable interconnections in the chiplet-based system. By integrating them into the silicon interposer, the 2T p-channel flash cells can potentially lead to the dense and flexible interconnection between multiple computing-in-memory (CIM) chiplets, resulting in highly reconfigurable and scalable chiplet-based CIM accelerators. A 180nm 1Kb 2T p-channel flash cell array is fabricated and characterized. The characterization results show the 2T p-channel flash cells exhibit a signal ratio >103over 1000 program/erase (P/E) cycles and the device-to-device variations are less than 21.07%. Their typical behaviors as routers are also confirmed by circuit simulations.
Weizeng Li, Linfang Wang, Zhi Li 0062, Wang Ye, Zhidao Zhou, Haiyang Zhou, Hanghang Gao, Jinshan Yue, Hongyang Hu, Fengman Liu, Chunmeng Dou
ISCAS12
2024 Write-Verify-Free MLC RRAM Using Nonbinary Encoding for AI Weight Storage at the Edge
abstract
High-density and reliable multilevel-cell (MLC) resistive random access memory (RRAM) is expected to meet the ever-increasing demand for on-chip weight storages in the intelligent edge devices. However, due to the device variations, many write-and-verify (WAV) iterations are usually required to program the RRAM cell, which causes high power consumption, long latency, and degradation on the memory lifetime. To address this issue, we propose a write–verify-free MLC RRAM macro for weight storage with 1) a cascode-current-mirror multibit write (CCM-MW) driver and 2) a nonbinary programming scheme (NB-PS) with a radix not greater than 2. A 180-nm 400-Kb RRAM test chip is demonstrated in silicon. For 2-bit-per-cell MLC storage, the value error rates can be reduced by 24.13% after introducing two redundant bits (RBDs). In addition, compared to the single-level cell (SLC) storage scheme, a 37.50% reduction in the number of cells can be achieved to store the ResNet-8 model with a 0.79% loss in inference accuracy without the need for WAV iterations.
Junjie An, Zhidao Zhou, Linfang Wang, Wang Ye, Weizeng Li, Hanghang Gao, Zhi Li 0062, Jinghui Tian, Hongyang Hu, Jinshan Yue, Lingyan Fan, Shibing Long, Qi Liu 0010, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.15
2023 A 40-nm SONOS Digital CIM Using Simplified LUT Multiplier and Continuous Sample-Hold Sense Amplifier for AI Edge Inference
abstract
Digital computing in memory (CIM) exhibits high precision as well as high energy efficiency (EE) yet still lacks discussion in nonvolatile memory (NVM). In this article, we propose a 40-nm silicon-oxide-nitride-oxide-silicon (SONOS)-based digital NVM CIM macro (DNV-CIM) featuring: 1) a simplified lookup table multiplier (SLUTM) combined with a lookup table (LUT) mapping scheme to improve area and EE and 2) a continuous sample-hold sense amplifier (CSH-SA) with an optimized voltage clamper and comparator for continuous read to reduce overall energy and time consumption for deep neural network (DNN) inference tasks. Performance evaluations indicate that the proposed DNV-CIM can achieve 93.04% accuracy and an EE up to 39.9 TOPS/W when running a 4-bit quantized ResNet18 trained on the CIFAR-10 dataset. This work presents a highly efficient digital CIM solution that can be readily implemented with commodity NVM.
Hongyang Hu, Haiyang Zhou, Danian Dong, Jinshan Yue, Wan Pang, Xiaoxin Xu, Chunmeng Dou
IEEE Trans. Very Large Scale Integr. Syst.9
2021 Sparsity-Aware Clamping Readout Scheme for High Parallelism and Low Power Nonvolatile Computing-in-Memory Based on Resistive Memory
abstract
The input parallelism of resistive memory (RRAM) based nonvolatile computing-in-memory (nvCIM) structure is limited by the signal margin as well as the readout precision. In this work, we propose a sparsity-aware clamping (SAC) scheme and its circuit implementation for nvCIM by co-design of circuit and algorithm. It can adaptively tune the quantized range and resolution of the readout circuit according to the degree of sparsity in neural network models. As a result, the SAC scheme can effectively increase the input parallelism of nvCIMs without incurring degradation on the signal margin or increasing the hardware cost for analogue readout. A case study on processing a multi-layer perceptron (MLP) model with the proposed nvCIM structure shows that the SAC scheme can improve the throughput by 2 times and increase the energy efficiency by 25.35% with negligible inference accuracy loss.
Linfang Wang, Wang Ye, Junjie An, Chunmeng Dou, Qi Liu 0010, Meng-Fan Chang, Ming Liu 0022
ISCAS4
2021 A 40nm 1Mb 35.6 TOPS/W MLC NOR-Flash Based Computation-in-Memory Structure for Machine Learning
abstract
Computation-in-memory (CIM) is a feasible method to overcome "Von-Neumann bottleneck" with high throughput and energy efficiency. In this paper, we proposed a 1Mb Multi-Level (MLC) NOR Flash based CIM (MLFlash- CIM) structure with 40nm technology node. A multi-bit readout circuit was proposed to realize adaptive quantization, which comprises a current interface circuit, a multi-level analog shift amplifier (AS-Amp) and an 8-bit SAR-ADC. When applied to a modified VGG-16 Network with 16 layers, the proposed MLFlash-CIM can achieve 92.73% inference accuracy under CIFAR-10 dataset. This CIM structure also achieved a peak throughput of 3.277 TOPS and an energy efficiency of 35.6 TOPS/W with 4-bit multiplication and accumulation (MAC) operations.
Sitao Zeng, Zhiguo Zhu, Zhaolong Qin, Chen Wang 0130, Jingjing Li 0001, Sanfeng Zhang 0001, Yajuan He, Chunmeng Dou, Xin Si, Meng-Fan Chang, Qiang Li 0021
ISCAS9