Meng Wu 0005

dblp:19/5921-5 · DBLP profile ↗
← Back
13ranked-venue papers
0as first author
13since 2021 · last 2026
0000-0002-7676-343XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 12 · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Quartet: A Digital Compute-in-Memory Versatile AI Accelerator With Heterogeneous Tensor Engines and Off-Chip-Less Dataflow
abstract
Although most AI core operations can be formulated as matrix multiplications (MMs), their characteristics are quite different. In some cases, one input may be a constant weight matrix while the other is a dynamic feature matrix, i.e., FWMM, or both inputs may be dynamic features as FFMM. Furthermore, the data involved can also be sparse, leading to variations like SpFWMM and SpFFMM. To address this challenge, this paper investigates a versatile accelerator architecture for AI algorithms based on heterogeneous tensor engines. For MM operators with varying characteristics, this paper proposes four-quadrant heterogeneous tensor engines to handle FWMM, SpFWMM, FFMM, and SpFFMM, respectively. These four tensor engines are comprised of SRAM-based single address digital compute-in-memory (CIM) array, SRAM-based multi-address digital CIM array, systolic array, and multi-SIMD array, respectively. In addition, to improve the AI execution efficiency, this paper proposes a dual-level multi-issue mechanism to achieve inter-operator and inter-block parallelization along with an off-chip-less dataflow enabled by on-chip unified memory pool. Thanks to the integration of aforementioned innovations, this paper develops a versatile AI acceleration chip Quartet, which achieves exceptional energy efficiency. Specifically, for graph convolutional network on PubMed, it demonstrates a$19.56\times $and$3.47\times $improvement in energy efficiency compared to similar works, ReDCIM and TensorCIM, respectively.
Yikan Qiu, Guoxiang Li, Meng Wu 0005, Yifan Jia 0009, Le Ye, Yufei Ma 0002
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 3D-TokSIM: Stacking 3D Memory with Token-Stationary Compute-in-Memory for Speculative LLM Inference
abstract
The LLM decoding process poses a significant challenge for memory bandwidth due to its autoregressive nature. Prior 2D memory solutions fail to overcome this memory bottleneck due to limited memory-to-logic bandwidth. In this work, we propose 3D-TokSIM, a cross-stack solution by stacking 3D memory on logic die with a specially designed token-stationary compute-in-memory (CIM) to efficiently accelerate speculative decoding. Our CIM is developed with novel token-stationary dataflow to reduce data movement on logic die to save power and balance computation and memory access. To further reduce the buffer requirements, we perform architecture exploration and allocate notable CIM resources for achieving higher decoding parallelism. Compared to RTX 3090 GPU, 3D-TokSIM achieves 15.1 $\times$ throughput and $324 \times$ energy efficiency improvements on speculative Llama2-7B decoding.
Boya Lv, Meng Wu 0005, Fengyun Yan, Yufei Ma 0002, Ru Huang 0001, Le Ye
DAC3
2025 Leveraging Compute-in-Memory for Efficient Generative Model Inference in TPUs
abstract
With the rapid advent of generative models, efficiently deploying these models on specialized hardware has become critical. Tensor Processing Units (TPUs) are designed to accelerate AI workloads, but their high power consumption neces-sitates innovations for improving efficiency. Compute-in-memory (CIM) has emerged as a promising paradigm with superior area and energy efficiency. In this work, we present a TPU architecture that integrates digital CIM to replace conventional digital systolic arrays in matrix multiply units (MXUs). We first establish a CIM-based TPU architecture model and simulator to evaluate the benefits of CIM for diverse generative model inference. Building upon the observed design insights, we further explore various CIM-based TPU architectural design choices. Up to 44.2% and 33.8% performance improvement for large language model and diffusion transformer inference, and 27.3 × reduction in MXU energy consumption can be achieved with different design choices, compared to the baseline TPUv4i architecture.
Zhantong Zhu, Hongou Li, Wenjie Ren, Meng Wu 0005, Le Ye, Ru Huang 0001
DATE4
2025 CIMTester: An Agile Golden-Result-Free BIST Compiler for Robust Compute-In-Memory
abstract
Digital compute-in-memory (DCIM) is playing an increasingly vital role in efficient AI computing due to its significant efficiency advantages. However, the combination of memory and computation logics in DCIM presents challenges for testing, including extra coupling fault between memory and logic, high overhead for golden result generation and indirect fault location. In this paper, we present a golden-result-free BIST structure together with a computation-coupled CIM BIST algorithm to reduce testing overhead and improve test coverage. In addition, we present CIMTester, a DCIM BIST compiler to adapt to swiftly changing test requirements and DCIM macro sizes and numbers in chips. The template-based generator generates BIST RTL based on proposed BIST structure and modifies the templates according to the architecture parameters of test chip. CIMTester’s iterator analyzes diverse sharing strategy of BIST components to satisfy user specifications of area, test time and fault coverage. We implemented and evaluated a series of TSMC 22nm DCIM macros with the generated BIST circuits, which achieves up to 99.48% fault coverage with only less than 2.44% area overhead.
Wenjie Ren, Meng Wu 0005, Le Ye
ICCAD2
2024 An In-Memory Computing Accelerator with Reconfigurable Dataflow for Multi-Scale Vision Transformer with Hybrid Topology
abstract
Transformer models equipped with multi-head attention (MHA) mechanism have demonstrated promise in computer vision (CV) tasks, i.e., vision transformers (ViTs). Nevertheless, the lack of inductive bias in ViTs leads to substantial computational and storage requirements, hindering their deployment on resource-constrained edge devices. To this end, multi-scale hybrid models are proposed to take the advantages of both transformers and convolutional neural networks (CNNs). However, existing domain-specific architectures focus on the optimization of either convolution or MHA at the expense of flexibility. In this work, an in-memory computing (IMC) accelerator is proposed to efficiently accelerate ViTs with hybrid MHA and convolution topology by introducing pipeline reordering. SRAM-based digital IMC macro is utilized to mitigate memory access bottleneck, while avoiding analog non-ideality. The reconfigurable processing engines and interconnections are investigated to enable the adaptable mapping of both convolution and MHA. Under typical workloads, experimental results exhibit that our proposed IMC architecture delivers 2.20× to 2.52× speedup and 40.6% to 74.8% energy reduction compared with the baseline design.
Zhiyuan Chen 0009, Yufei Ma 0002, Yifan Jia 0009, Guoxiang Li, Meng Wu 0005, Le Ye, Ru Huang 0001
DAC6
2024 AIG-CIM: A Scalable Chiplet Module with Tri-Gear Heterogeneous Compute-in-Memory for Diffusion Acceleration
abstract
The emergence of Diffusion models has gained significant attention in the field of Artificial Intelligence Generated Content. While Diffusion demonstrates impressive image generation capability, it faces hardware deployment challenges due to its unique model architecture and computation requirement. In this paper, we present a hardware accelerator design, i.e. AIG-CIM, which incorporates tri-gear heterogeneous digital compute-in-memory to address the flexible data reuse demands in Diffusion models. Our framework offers a collaborative design methodology for large generative models from the computational circuit-level to the multi-chip-module system-level. We implemented and evaluated the AIG-CIM accelerator using TSMC 22nm technology. For several Diffusion inferences, scalable AIG-CIM chiplets achieve 21.3× latency reduction, up to 231.2× throughput improvement and three orders of magnitude energy efficiency improvement compared to RTX 3090 GPU.
Yiqi Jing, Meng Wu 0005, Yufei Ma 0002, Ru Huang 0001, Le Ye
DAC2
2024 DCIM-GCN: Digital Computing-in-Memory Accelerator for Graph Convolutional Network
abstract
Graph convolutional network (GCN) has gained great success in a diverse range of intelligent tasks. However, the hardware performance of GCNs is often bounded by random and non-continuous memory accesses due to the sparse graph data, which incur high latency and high power consumption. The emerging computing-in-memory (CIM) architecture significantly reduces the overhead of data movements, which is suitable for memory-intensive GCN acceleration. Existing analog-based CIM solutions require a large amount of analog-to-digital (AD) and digital-to-analog (DA) conversions, which dominate the overall area and power consumption. Furthermore, the analog non-ideality can degrade accuracy and reliability of CIM. To address these challenges, this work proposes a digital CIM accelerator based on SRAM, called DCIM-GCN, to accelerate GCN algorithm. DCIM-GCN introduces innovations on three levels: circuit, architecture, and algorithm. At the circuit level, digital CIM is proposed with SRAM sub-arrays to eliminate the power and area expensive AD/DA converters. Furthermore, we have incorporated the multi-address feature into the digital CIM, thereby leveraging its ability to efficiently process sparse matrix multiplication. At the architecture level, the sparsity-aware computation engine takes advantage of sparsity in GCNs and leverages CIM to minimize memory accesses and data movements. Finally, at the algorithm level, the balance mapping algorithm tackles workload imbalance issues, while the vertex reorder algorithm reduces idle states for aggregation engines, resulting in increased hardware utilization. Our DCIM-GCN achieves 1.89$\times$and 2.42$\times$speedup and 4.58$\times$and 9.46$\times$energy efficiency improvement on average over other CIM-based graph accelerators, e.g., PASGCN and PIM-GCN, respectively.
Yufei Ma 0002, Yikan Qiu, Guoxiang Li, Meng Wu 0005, Le Ye, Ru Huang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.5
2023 RIMAC: An Array-Level ADC/DAC-Free ReRAM-Based in-Memory DNN Processor with Analog Cache and Computation
abstract
By directly computing in analog domain, processing-in-memory (PIM) is emerging as a promising alternative to overcome the memory bottleneck of traditional von-Neuman architecture, especially for deep neural networks (DNNs). However, the data outside PIM macros in most existing PIM accelerators are stored and operated as digital signals that require massive expensive digital-to-analog (D/A) and analog-to-digital (A/D) converters. In this work, an array-level ADC/DAC-free ReRAM-based in-memory DNN processor named RIMAC is proposed, which accelerates various DNNs in pure analog-domain with analog cache and analog computation modules to eliminate the expensive D/A and A/D conversions. Our experiment result shows the peak energy efficiency is improved by about 34.8×, 97.6×, 10.7×, and 14.0× compared to PRIME, ISAAC, Lattice, and 21'DAC for various DNNs on ImageNet, respectively.
Meng Wu 0005, Yufei Ma 0002, Le Ye, Ru Huang 0001
ASP-DAC2
2023 DCIM-3DRec: A 3D Reconstruction Accelerator with Digital Computing-in-Memory and Octree-Based Scheduler
abstract
Learning-based 3D reconstruction has evolved rapidly with promising quality, while it requires high-performance hardware for interactive applications. In this work, a reconstruction accelerator called DCIM-3DRec is presented which leverages digital computing-in-memory (DCIM) design to facilitate learning-based reconstruction deployment on realtime and low-power edge platforms. The DCIM-3DRec is designed with the following features: a reconfigurable DCIM macro array for high data reuse and macro utilization, and an Octree-based subdivision scheduler for efficient management of 3D space prediction. The DCIM-3DRec accelerator is implemented and evaluated in TSMC 55 nm technology, with a DCIM macro efficiency of 19.4 TOPS/W at INT8. Overall, the DCIM-3DRec accelerator achieves 23× performance gain and four orders of magnitude energy efficiency improvement compared to a Nvidia RTX3090 GPU.
Yiqi Jing, Meng Wu 0005, Fengyun Yan, Yufei Ma 0002, Le Ye
ISLPED5
2023 Research progress on low-power artificial intelligence of things (AIoT) chip design
Le Ye, Zhixuan Wang, Yufei Ma 0002, Linxiao Shen, Yihan Zhang 0002, Meng Wu 0005, Ying Liu 0069, Yiqi Jing, Hao Zhang 0119, Ru Huang 0001
Sci. China Inf. Sci.9
2022 DCIM-GCN: Digital Computing-in-Memory to Efficiently Accelerate Graph Convolutional Networks
abstract
Computing-in-memory (CIM) is emerging as a promising architecture to accelerate graph convolutional networks (GCNs) normally bounded by redundant and irregular memory transactions. Current analog based CIM requires frequent analog and digital conversions (AD/DA) that dominate the overall area and power consumption. Furthermore, the analog non-ideality degrades the accuracy and reliability of CIM. In this work, an SRAM based digital CIM system is proposed to accelerate memory intensive GCNs, namely DCIM-GCN, which covers innovations from CIM circuit level eliminating costly AD/DA converters to architecture level addressing irregularity and sparsity of graph data. DCIM-GCN achieves 2.07X, 1.76X, and 1.89× speedup and 29.98×, 1.29×, and 3.73× energy efficiency improvement on average over CIM based PIMGCN, TARe, and PIM-GCN, respectively.
Yikan Qiu, Yufei Ma 0002, Meng Wu 0005, Le Ye, Ru Huang 0001
ICCAD4
2022 Reliability-Improved Read Circuit and Self-Terminating Write Circuit for STT-MRAM in 16 nm FinFET
abstract
High power consumption is usually required in a spin-torque-transfer magnetoresistive random access memory (STT-MRAM) array’s peripheral circuits for reliable operations. In read, power needs to be spent for the low absolute resistance in the magnetic tunnel junctions (MTJ), and a limited high-state-to-low-state resistance ratio calls for high currents for the same detectable readout voltage under accuracy requirements. In write, the random programming time poses challenges for energy efficient write operations within an acceptable write error rate. To address the issues mentioned above, in this work, we propose a reliability-improved read circuit that consumes only 92.09 fJ/bit read energy while ensuring correct readout values under 4.5 sigma resistance variance, and a self-terminating write peripheral circuit achieving an energy reduction of 82.3% at 1 part-per-million write error rate (WER) under 20 ns write period.
Chang Xue, Yihan Zhang 0002, Mingwei Zhu, Tianqiao Wu, Meng Wu 0005, Yandong He, Le Ye
ISCAS6
2021 The Challenges and Emerging Technologies for Low-Power Artificial Intelligence IoT Systems
abstract
The Internet of Things (IoT) is an interface with the physical world that usually operates in random-sparse-event (RSE) scenarios. This article discusses main challenges of IoT chips: power consumption, power supply, artificial intelligence (AI), small-signal acquisition, and evaluation criteria. To overcome these challenges, many works recently aimed at IoT system design have emerged. This work reviews the architecture and circuit innovations that have contributed to IoT developments. This paper does not cover security of IoT. Event-driven architectures and nonuniform sampling ADCs significantly reduce the long-term average power. Besides, embedding AI engines in IoT nodes (AIoT) is one critical trend. The computing-in-memory technique improves the energy efficiency of the AI engine. Asynchronous spike neural networks (ASNNs) AI engines show low power potential. In addition to data processing, small-signal acquisition is also critical. The charge-domain analog-front-end (AFE) techniques such as floating inverter-based amplifiers improve energy efficiency. In addition to the above low power and high energy efficiency technologies, energy harvesting can also enhance the lifetime of AIoT devices. This article discusses recent ambient RF and natural energy harvesting approaches and high-efficiency DC-DC with a wide load range. Finally, novel evaluation criteria are introduced to establish benchmark standards for AIoT chips.
Le Ye, Zhixuan Wang, Ying Liu 0069, Hao Zhang 0119, Meng Wu 0005, Linxiao Shen, Yihan Zhang 0002, Zhichao Tan, Yangyuan Wang, Ru Huang 0001
IEEE Trans. Circuits Syst. I Regul. Pap.7