Zongwei Wang 0001

dblp:125/8211 · DBLP profile ↗
← Back
26ranked-venue papers
0as first author
21since 2021 · last 2026
0000-0001-6297-2700ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 17 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 3 since 2021Software engineering, systems software and programming languages · 4 · 4 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CHIP-MAP: A Collaborative Optimization Framework for Macro Placement Using Large Language Models
abstract
As integrated circuits continue to grow in both scale and complexity, macro placement plays a critical role in physical design, directly affecting chip-level performance, power, and area (PPA). Traditional macro placement methods, such as simulated annealing, analytical optimization, and reinforcement learning, face limitations including slow convergence, heavy dependence on large datasets, and over-reliance on intermediate PPA indicators rather than final PPA. Large language models (LLMs) offer strong generative power and semantic reasoning that can potentially automate macro layout tasks while addressing the aforementioned problems in traditional methods, but their limited understanding of layout rules and lack of iterative, feedback-driven refinement make direct application challenging. To address this, we propose CHIP-MAP, a macro placement framework based on multi-agent collaboration and feedback-driven optimization. Furthermore, we introduce two innovative tools: the Module Link Weight Analyzer (MWA) and the Standard Cell Usability Score (SCUS), which are designed to guide fine-grained layout refinement. We evaluate CHIP-MAP on five benchmarks ranging from low-power cores to large multi-core processors implemented at 130nm and 45nm technology nodes. Results show that it achieves up to 1.5% area reduction and an average repair of 61.6% of total negative slack (TNS), while also reducing wirelength and improving timing.
Yiming Du, Renye Yan, Yunfan Yang, Frank Qu, Jiajun Tan, ZhiYu Zheng, Yiming Gan, Ling Liang 0003, Zongwei Wang 0001, Yimao Cai
DATE9
2026 GMaC: NvCIM Architecture for Parallel Point-based Point Cloud Acceleration via Geometric Mapping and Address-Index Computation
Zongwei Wang 0001, Ling Liang 0003, Yimao Cai
DATE2
2026 SONIC: Smart Optimization for Neural-Integrated CMP with Timing-Aware Fills
abstract
Dummy fill insertion is essential for CMP uniformity but remains challenging due to the nonlinear CMP process, the large optimization space, and timing degradation caused by parasitic coupling. We propose SONIC, a differentiable CMP-driven dummy fill optimization framework that employs a neural CMP simulator to directly optimize planarization objectives using gradient-based methods. SONIC further integrates a timing-aware fill insertion strategy to mitigate coupling capacitance near critical nets. Experimental results demonstrate that SONIC achieves competitive planarization quality with up to 1830× runtime speedup over a full-chip CMP simulator. Compared with the state-of-the-art model-based method, SONIC reduces height variation, line deviation, and outliers by up to 86.16%, 90.10%, and 51.61%, respectively, while achieving a 77.67% runtime reduction and lowering coupling capacitance by 13.05%.
Jiajun Tan, Yiming Du, Yiming Gan, Ling Lang 0002, Yibo Lin, Zongwei Wang 0001, Yimao Cai
DATE8
2026 DIAMoND: Dynamic Inference for Adaptive Edge MOE with Heterogeneous In-NAND and Near-DRAM Compute Architecture
Tianyang Luo, Shuzhang Zhong, Dongxue Zhao, Renjie Wei, Meng Li 0004, Guangyu Sun 0003, Zongwei Wang 0001, Yimao Cai
ISCA10
2026 A Sparsity-Aware Reconfigurable Sensing-Quantization Circuit for RRAM-Based Analog Compute-in-Memory
Zhuoya Chen, Hao Ding 0011, Yunfan Yang, Haisu Zhang, Jinshan Li, Shigeng Zhao, Yunyi Fu, Zongwei Wang 0001, Yimao Cai
ISCAS8
2026 An IZO-Based 2T0C Compute-in-Memory Array with Adaptive Read Voltage Boosting for Energy-Efficient Edge AI
Hao Ding 0011, Jiye Li, Xiantong Qiu, Yunfan Yang, Gaoqi Yang, Shengdong Zhang, Zongwei Wang 0001, Yimao Cai
ISCAS9
2026 A Near-Sensor Image Compression Architecture with RRAM-based Hyperdimensional Encoder for Smart Vision
Haoyang Gu, Zongwei Wang 0001, Jingshan Li, Zezhi Chen, Ling Liang 0003, Wengao Lu, Yimao Cai
ISCAS3
2026 RRAM-based CAM for Energy-Efficient In-Memory Text Compression System
Lianliang Wu, Hao Ai, Ao Shi, Haokai Guan, Kexun Li, Kefan Tao, Yulin Feng, Zongwei Wang 0001, Yimao Cai, Peng Huang 0004
ISCAS12
2026 A 1.8-ns, 45fJ/bit Time-Domain Sensing Scheme with Offset-Cancelled Resistance-to-Time Converter and 10-5 BER for Digital RRAM Compute-in-Memory
Shigeng Zhao, Hao Ding 0011, Yunfan Yang, Jinshan Li, Zhuoya Chen, Xing Zhang 0002, Zongwei Wang 0001, Yimao Cai
ISCAS7
2026 Efficient LoRA-Based Weight Update Write-Back in 3D NAND Flash for Large Language Models
Dongxue Zhao, Tianyang Luo, Ling Liang 0003, Zongwei Wang 0001, Yimao Cai
ISCAS4
2026 OmniGuard: Two-Level Protection Framework for RRAM-Based Accelerator With High Efficiency and Flexibility
abstract
RRAM-based Deep Neural Network (DNN) accelerators have gained widespread usage in edge devices. However, the security vulnerabilities of RRAM-based accelerators hinder their real application. Current research on safeguarding RRAM-based accelerators predominantly relies on a single-level protection approach. This has resulted in the restriction of its protection scope, the rigidity and lack of generality in the protection method, or has had an impact on the computational efficiency of the system. As a result, it encounters substantial challenges in attaining comprehensive optimization across multiple dimensions, such as universality, the scope of protection, and security-related overheads. In this paper, we develop specific analyses on accelerators and attacks and build graph-based representations. Based on these, we partition the RRAM-based accelerators’ security into two levels: on-chip security and off-chip security. Furthermore, we propose a two-level protection framework for RRAM-based accelerators, which is calledOmniGuard. At the on-chip security level,OmniGuardproposes a bit-grained shuffle to achieve protection while using lightweight Benes Networks to maintain the CIM capability. At the off-chip level,OmniGuardproposes an RRAM-based AES engine to introduce the AES algorithm into the accelerator with significant acceleration and minimal overhead. Evaluation results demonstrate thatOmniGuardprovides powerful and flexible protection while achieving 1.33×∼4.38× speedup and 1.26×∼2.65× power savings, with only 5% energy overhead and 3% area overhead.
Ling Liang 0003, Yunfan Yang, Jinlong Lin, Meng Li 0004, Zongwei Wang 0001, Yimao Cai
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.6
2026 REF-CIM: A 40-nm Non-Ideality Tolerant and Energy Efficient RRAM Compute-in-Memory Macro With Configurable Precision for Edge AI
Hao Ding 0011, Yunfan Yang, Zongwei Wang 0001, Jinshan Li, Lin Bao, Ling Liang 0003, Yimao Cai
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 SA-CIM: A 28nm 16Mb RRAM-based Sparsity-Aware Compute-In-Memory Macro for Edge AI Algorithm Processing
abstract
Compute-in-memory (CIM) for edge devices is usually constrained by on-chip resources, including on-chip memory and physical chip size, which hinders the deployment of more complex neural networks. By leveraging the sparsity of neural networks, the overall memory requirements and energy consumption can be reduced. However, Existing sparsity-aware architectures cannot achieve high energy efficiency due to off-chip sparsity control. This work proposes:1) Hybrid sparsity regulation strategy. The sparsity encoding and alignment circuit is designed and implemented, realizing on-chip sparsity detecting and encoding. 2) Sparsity-aware compute-in-memory (CIM) array based on RRAMs. The in-situ deployment of unstructured sparsity is implemented inside the CIM array, and the CIM array and sparsity are tightly coupled by sparsity read/write. This work demonstrates the design and evaluation of SA-CIM: a sparsity-aware CIM macro with 16Mb RRAM with fine-grained sparsity detecting and encoding capacity, achieving energy efficiency of 22.7TOP/W@8b/8b.
Hao Ding 0011, Zongwei Wang 0001, Jinshan Li, Shigeng Zhao, Heting Gao, Junbo Ao, Ling Liang 0003, Yimao Cai, Ru Huang 0001
ISCAS2
2025 HRC-CIM: Hybrid RRAM-Capacitor Cell based Compute-in-Memory with High Linearity, Parallelism and Energy Efficiency
abstract
RRAM-based Compute-in-memory (CIM) has emerged as a promising computing paradigm for artificial intelligence (AI) algorithms. However, the low on/off ratio and high on-current have been the major challenges to enhance the accuracy, parallelism, and energy efficiency. In this paper, we propose a novel Hybrid RRAM-Capacitor (HRC) cell based CIM macro to address these issues. The proposed HRC cell achieves a high on-off ratio with sub-100nA on-current and eliminates direct current path during computation, which significantly enhances both parallelism and energy efficiency. The write-verify scheme for RRAM programming is optimized for HRC cell array, and is further supported by a quantization result calibration technique using a dummy column to ensure high linearity and accuracy in analog domain multiply-and-accumulate (MAC) operations. A HRC-CIM macro has been designed and demonstrated using 28nm technology node, enabling block-level parallelism across 64 rows with 4-bit input per row, and delivering an energy efficiency of up to 40.40 TOPS/W @8b-IN/8b-W.
Jinshan Li, Zongwei Wang 0001, Hao Ding 0011, Yunfan Yang, Shigeng Zhao, Shengyu Bao, Ruiqing Xie, Zhuoya Chen, Yimao Cai, Ru Huang 0001
ISCAS2
2025 Polynomial geometric transformation based on IGZO charge trapping RAM array for machine vision calibration
Lin Bao, Haisu Zhang, Zongwei Wang 0001, Linbo Shan, Cuimei Wang, Yimao Cai, Shanguo Huang
Sci. China Inf. Sci.3
2025 Matrix: Multi-Cipher Structures Dataflow for Parallel and Pipelined TFHE Accelerator
abstract
Fully homomorphic encryption over torus (TFHE) enables the execution of arbitrary functions on encrypted data through programmable bootstrapping (PBS). However, performing all operations on ciphertext during PBS results in high computational and memory requirements, limiting the deployment of PBS in real-world scenarios. Previous TFHE accelerator designs have attempted to improve performance by employing specific dataflow and functional units, but these techniques may require large off-chip bandwidth or on-chip storage when scaling up computation capacity. Additionally, the design of specialized functional units may limit the utilization of computation units when facing dynamic secure parameter settings. To address these challenges and further improve PBS throughput in TFHE, we propose Matrix , an ASIC-based architecture that balances off-chip bandwidth and on-chip storage according to the execution flow of PBS. In Matrix , we utilize a unified special-prime-based processing element (PE) that achieves high utilization with minimal resource overhead. Furthermore, we propose a hybrid PBS dataflow that can efficiently reduce computation complexity and memory requirements. Compared to state-of-the-art TFHE accelerators, Matrix achieves 1.43 × -5.66 × throughput improvement for PBS. For ZAMA Deep-NN benchmark, we achieve 525.60× and 68.06× speedup compared to CPU and GPU, respectively. 1
Ling Liang 0003, Fahong Zhang 0004, Zhirui Li, Xin Fan 0009, Dimin Niu, Meng Li 0004, Zhiyong Li 0016, Zongwei Wang 0001, Hongzhong Zheng, Yimao Cai, Yuan Xie 0001
ACM Trans. Archit. Code Optim.10
2024 An isolated symmetrical 2T2R cell enabling high precision and high density for RRAM-based in-memory computing
Yaotian Ling, Zongwei Wang 0001, Lin Bao, Shengyu Bao, Yimao Cai, Ru Huang 0001
Sci. China Inf. Sci.2
2024 Investigation and mitigation of Mott neuronal oscillation fluctuation in spiking neural network
Lindong Wu, Zongwei Wang 0001, Lin Bao, Linbo Shan, Zhizhen Yu, Yunfan Yang, Shuangjie Zhang, Guandong Bai, Cuimei Wang, John Robertson, Yuan Wang 0001, Yimao Cai, Ru Huang 0001
Sci. China Inf. Sci.2
2023 Mnemonic Dictionary Learning for Intrinsic Motivation in Reinforcement Learning
abstract
Reinforcement learning for hard-exploration tasks remains challenging due to the long-term dependence and sparse-and-delay rewards in complex environments. In these challenging tasks, intrinsic motivation has become a dominant paradigm to enable the agent to explore the environment when no external reward feedback is available. In this work, inspired by studies from the human memory mechanism, we present a mnemonic dictionary learning (MDL) model for intrinsic motivation in reinforcement learning. The MDL model leverages sparse dictionary learning to incremental abstract the exploration histories into a compact memory-like dictionary, providing an excellent intrinsic motivation model. This mnemonic dictionary model not only drives the agent to explore novel stats in the environments indicated by the memory reconstruction error but also helps the agent to remember the key states and structure of the environments using its learned bases and reconstruction coefficients. The proposed MDL model can serve as a generative module for existing exploration methods. Extensive experimental results on typical sparse-reward tasks demonstrate its effectiveness and applicability over several competing algorithms. We will release the source code and trained models to facilitate further studies in this research direction.
Renye Yan, Yuan Zhan, Pin Tao, Zongwei Wang 0001, Yimao Cai, Junliang Xing
IJCNN5
2022 PIMulator-NN: An Event-Driven, Cross-Level Simulation Framework for Processing-In-Memory-Based Neural Network Accelerators
abstract
Processing-in-memory (PIM) architecture has been proposed to accelerate state-of-the-art neuro-inspired algorithms, such as deep neural networks. In this article, we present PIMulator-NN, an event-driven, cross-level simulation framework for PIM-based neural network accelerators. By employing an event-driven simulation mechanism, PIMulator-NN is able to model architecture details and capture design details of the architecture. Moreover, we integrate the main-stream circuit-level simulation framework with PIMulator-NN to accurately simulate the area, latency, and energy consumption of analog computation units. To demonstrate the usage of PIMulator-NN, we implement several PIM designs with PIMulator-NN and perform detailed simulation. The simulation results show that memory access and interconnects make considerable impacts on system-level performance and energy. Note that such results are hard to be captured by conventional performance model-based estimations. We found some anti common-sense results while modeling the architecture details with PIMulator-NN. With several architecture templates, PIMulator-NN provides the users with a platform to build up their PIM architecture quickly. PIMulator-NN is able to capture the impacts of different design choices (e.g., dataflow, interconnect, data parallelism, etc.), and this could enable users to explore their design space efficiently.
Qilin Zheng, Yijin Guan, Zongwei Wang 0001, Yimao Cai, Yiran Chen 0001, Guangyu Sun 0003, Ru Huang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2021 A High Accuracy Multiple-Command Speech Recognition ASIC Based on Configurable One-Dimension Convolutional Neural Network
abstract
Speech command interaction has drawn much attention in smart application market. Many of previous chips achieve an ultra-low power consumption at the cost of a certain accuracy loss, and they are designed only for the fixed speech command recognition tasks, which is inflexible and restrains further development. Here, we demonstrate a configurable speech command recognition ASIC with an ultra-high accuracy fabricated by the TSMC commercial 180-nm CMOS technology. In this chip, Mel-Frequency Cepstrum Coefficients (MFCCs) are used as speech features and a One-Dimension Convolutional Neural Network (1-D CNN) is adopted for the speech feature recognition, which simplifies the design of network and the storage method of memory. Moreover, the configurable 1-D CNN layer of the network ensures the diversity and flexibility of the commands. The measurement results indicate that the chip achieves a 95.6% accuracy on Google Speech Command Database (GSCD) when working at 16 MHz and keeping a reasonable power consumption as 26.4 mW. Moreover, the chip supports max 30 speech commands at a time, which is better than the state-of-the-art chips.
Lindong Wu, Zongwei Wang 0001, Yimao Cai, Ru Huang 0001
ISCAS2
2020 Lattice: An ADC/DAC-less ReRAM-based Processing-In-Memory Architecture for Accelerating Deep Convolution Neural Networks
abstract
Nonvolatile Processing-In-Memory (NVPIM) has demonstrated its great potential in accelerating Deep Convolution Neural Networks (DCNN). However, most of existing NVPIM designs require costly analog-digital conversions and often rely on excessive data copies or writes to achieve performance speedup. In this paper, we propose a new NVPIM architecture, namely, Lattice, which calculates the partial sum of the dot products between the feature map and weights of network layers in a CMOS peripheral circuit to eliminate the analog-digital conversions. Lattice also naturally offers an efficient data mapping scheme to align the data of the feature maps and the weights and hence, avoiding the excessive data copies or writes in the previous NVPIM designs. Finally, we develop a zero-flag encoding scheme to save the energy of processing zero-values in sparse DCNNs. Our experimental results show that Lattice improves the system energy efficiency by 4× ~ 13.22× compared to three state-of-the-art NVPIM designs: ISAAC, PipeLayer, and FloatPIM.
Qilin Zheng, Zongwei Wang 0001, Zishun Feng, Bonan Yan, Yimao Cai, Ru Huang 0001, Yiran Chen 0001, Chia-Lin Yang, Hai Li 0001
DAC2
2020 MobiLattice: A Depth-wise DCNN Accelerator with Hybrid Digital/Analog Nonvolatile Processing-In-Memory Block
abstract
Nonvolatile Processing-In-Memory (NVPIM) architecture is a promising technology to enable energy-efficient inference of Deep Convolutional Neural Networks (DCNNs). One major advantage of NVPIM is that the vector dot-product operations can be completed efficiently by analog computing inside a Nonvolatile Memory (NVM) crossbar. However, its inference efficiency is severely downgraded when processing depth-wise convolution layers, which have been widely employed in many lightweight DCNNs. One major challenge is that the cell utilization is extreme low when mapping the depth-wise convolution layer to a crossbar. To overcome this problem, we propose a novel hybrid mode NVPIM architecture, namely, MobiLattice. With moderate hardware overhead, MobiLattice enables both analog and digital mode operations on NVM crossbars. While conventional convolution layers are computed efficiently using the analog mode, the computation efficiency of depth-wise convolution layers are substantially improved using the digital mode by mitigating the redundant memory space in the NVM crossbars. Experimental results show that, compared to prior approaches where only the analog mode is supported by the NVPIM architecture, MobiLattice can speedup the processing of typical depth-wise DCNNs by 2 ~ 5× on average and up to 30× by combining with some extreme quantization schemes.
Qilin Zheng, Zongwei Wang 0001, Guangyu Sun 0003, Yimao Cai, Ru Huang 0001, Yiran Chen 0001, Hai Li 0001
ICCAD3
2019 Enhance the Robustness to Time Dependent Variability of ReRAM-Based Neuromorphic Computing Systems with Regularization and 2R Synapse
abstract
Time Dependent Variability (TDV) is one of the major concerns in implementing a Neuromorphic Computing System (NCS) with Resistive Random Access Memory (ReRAM). In this work, we propose a variation-distribution aware training algorithm to enhance the robustness of NCS to TDV without incurring extra hardware overhead by leveraging algorithm-level regularization and hardware-level 2R synapse structure. Simulation results on image recognition tasks show that our method improves the system accuracy by up to ∼4% and ∼10% under the worst-case TDV condition for MNIST and CIFAR-10, respectively. Detailed analysis also shows that our method allows the NCS to use synapses with higher resistance than conventional design for the same accuracy requirement, introducing potential energy saving.
Qilin Zheng, Zongwei Wang 0001, Yimao Cai, Ru Huang 0001, Bing Li 0017, Yiran Chen 0001, Hai Li 0001
ISCAS3
2019 Investigation of NbOx-based volatile switching device with self-rectifying characteristics
Yichen Fang, Zongwei Wang 0001, Caidie Cheng, Zhizhen Yu, Yuchao Yang 0001, Yimao Cai, Ru Huang 0001
Sci. China Inf. Sci.2
2018 Integration of biocompatible organic resistive memory and photoresistor for wearable image sensing application
Yichen Fang, Zongwei Wang 0001, Yuchao Yang 0001, Jintong Xu, Yimao Cai, Ru Huang 0001
Sci. China Inf. Sci.4