Sifan Sun

dblp:309/0985 · DBLP profile ↗
← Back
9ranked-venue papers
2as first author
9since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 8 · 1 first-author · 8 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 High-Efficiency and Low-Deviation Analog-Digital Hybrid Compute-in-Memory Architecture With Dynamic Weight Division
abstract
Compute-in-memory (CIM) reduces data movement but suffers from an accuracy–efficiency trade-off: Analog CIM (ACIM) is energy-efficient but loses accuracy and incurs higher cost at large bit-widths, while digital CIM (DCIM) supports high precision but is inefficient for low-precision tasks. To overcome these challenges, we propose an analog–digital hybrid CIM (HCIM) architecture to address this trade-off, including 1) an analog–digital hybrid 10T SRAM cell without additional transistors and a dual-capacitor-based multicycle weighting module to reduce area; 2) a successive-approximation-register (SAR) ADC with a pseudo C-2C capacitor array that can be reconfigured from an 8-bit ADC into two parallel 4-bit ADCs to improve configurability; 3) configurable weight division and computing resource allocation strategies. Simulations in a 28-nm process show that HCIM achieves 15.56 TOPS/W at 12-bit ($8+4$) with$1.33\times $and$2.35\times $efficiency improvement over DCIM and ACIM and$16\times $lower error. It achieves 27.87 TOPS/W at 8-bit and 78.13 TOPS/W at 4-bit, demonstrating superior energy efficiency, computational accuracy, and flexibility.
Linjun Jiang, Sifan Sun, Wente Yi, Dengwen Li, Wang Kang 0001, He Zhang 0011, Weisheng Zhao 0001
IEEE Trans. Circuits Syst. I Regul. Pap.3
2026 A 40 nm Buffer-Free 7T-SRAM Analog Charge-Domain CIM Macro With Merging Timing Based On Time-Row Division Strategy
abstract
Computing-in-memory (CIM) macros based on static random access memory (SRAM) are meant to increase capacity while improving energy efficiency and reducing computing latency. However, traditional analog designs still face several key challenges, including long computing latency from separated computing phases, negative voltage fluctuations from massive parallel computing, and low bitcell density from additional transistors and capacitors for multiplication. On the other hand, only time-aligned inputs are supported in the works. To overcome the above challenges, this work proposes a buffer-free 7T-SRAM charge-domain CIM macro. It has four key features: 1) a compact 7T SRAM bitcell structure for high-energy efficiency; 2) a configurable input unit to support different sizes of input activations; 3) a time-row division (RD) strategy to support real-time processing and alleviate negative voltage fluctuations; and 4) a merging timing to conceal the input phase for high throughput. The fabricated 512-Kb SRAM-CIM macro in 40 nm achieves 79.3–290.4 Tops/W at 4-bit precision.
Linjun Jiang, Sifan Sun, Changyu Li, Wang Kang 0001, He Zhang 0011
IEEE Trans. Very Large Scale Integr. Syst.3
2025 A 24.65 TOPS/W@INT8 Hybrid Analog-Digital Multi-core SRAM CIM Macro with Optimal Weight Dividing and Resource Allocation Strategies
abstract
Compute-in-memory (CIM) technology integrates memory and computation to reduce memory bottlenecks in modern systems. However, current CIM architectures face challenges in balancing accuracy and energy efficiency. Analog-CIM (ACIM) is energy-efficient but less accurate, while Digital-CIM (DCIM) is accurate but consumes more energy. In this paper, we propose a novel multi-core hybrid analog-digital CIM macro that effectively addresses this trade-off. Our approach intelligently allocates computation tasks to ACIM and DCIM cores based on their accuracy requirements, achieving a balance of accuracy and efficiency. Additionally, we developed an optimization framework to determine the optimal weight divide ratio and computing resource allocation for the hybrid CIM. Experimental results demonstrate the efficacy of our approach. The proposed hybrid CIM achieves an outstanding energy efficiency of 24.65 TOPS/W at 8-bit precision, surpassing DCIM by a factor of 1.33 while maintaining a low error rate of only 0.4%, which is 30 times better than ACIM at the same precision.
Wente Yi, Sifan Sun, Wenjia Wang 0011, Jinyu Bai, He Zhang 0011, Wang Kang 0001
ASP-DAC3
2025 Efficient Weight Mapping and Resource Scheduling on Crossbar-based Multi-core CIM Systems
abstract
Crossbar-based computing-in-memory (CIM) systems facilitate large-scale parallel multiply-and-accumulate (MAC) operations, while a domain-specific compiler (DSC) plays a pivotal role in optimizing the deployment of neural network algorithms on such systems. With the development of multi-core and large-core architectures, some key compiler problems such as high parallel processing, resource utilization, and crossbar array assignment methods have not been solved. For low-latency application scenarios, we have designed a resource scheduling strategy for our hardware system based on stream data processing to reduce the latency caused by intra-core and intercore communication. Additionally, a weight mapping strategy has been developed to maximize the potential of crossbar arrays in convolutional neural networks (CNNs) deployment. Experimental results on our multi-core eFlash-based CIM system-on-chip (SoC) demonstrate that these two technologies help CNNs achieve a 76% reduction in latency, a 30% improvement in resource utilization, and the use rate of crossbar array that can reach up to 94.7%.
Sifan Sun, Aifei Zhang, Haiyan Qin, Minhao Gu, Shihang Fu, Shuaikai Liu, Baosen Liu, Wang Kang 0001
DAC2
2025 Model quantization for computing-in-memory: a survey
Sifan Sun, Jinyu Bai, Hanting Chen, Kaiwen Deng, Zhiwei Xie 0009, He Zhang 0011, Wang Kang 0001, Weisheng Zhao 0001
Sci. China Inf. Sci.1
2025 A Self-Decryption Pass Transistor Logic-Based In-MRAM Computing Macro Using Hybrid VGSOT-MTJ/GAA-CNTFET
abstract
Spintronic devices and gate-all-around carbon nanotube field-effect-transistors (GAA-CNTFETs)-based computing in-memory architecture are competitive candidates for applications in battery-powered tiny artificial intelligence (AI) edge devices. Meanwhile, data encryption and decryption are also necessary to protect AI model weights and the customized data used to guarantee neural network (NN) inference accuracy. In this study, we propose a self-decryption pass transistor logic (PTL)-based in-MRAM computing macro (SP-CIM) that utilizes hybrid voltage-gated spin-orbit torque magnetic tunnel junctions (VGSOT-MTJ)/GAA-CNTFET. The proposed SP-CIM macro enables simultaneous data access, decryption, and full-accuracy multiply-and-accumulate (MAC) operations using the newly introduced voltage-divider self-decryption cell, without the need for additional decryption logic. Compared to existing in-memory decryption strategies, this design reduces energy consumption by 45.7% and decreases decryption delay by 87.2%. To enhance area efficiency and reduce computing latency, we propose a PTL-based multiplication cell that achieves full-accuracy local 2b-IN TEXPRESERVE0 2b-W operations with only 20 transistors (20T). Additionally, novel PTL-based full-swing output half adders (10T-HA) and full adders (14T-FA) are proposed to construct the local adder tree, achieving reductions of 31.8%, 76.4%, and 41.4% in energy, delay, and area, respectively, compared to conventional adder trees in CIM macros. Simulations of the 288 kb SP-CIM macro demonstrated throughput and energy efficiency of 2.25 TOPS and 226.6 TOPS/W, respectively, at a 0.6 V supply voltage, and 2.97 TOPS and 154.1 TOPS/W, respectively, at a 0.8 V supply voltage, with 8b-IN, 8b-W, and 24b-OUT.
Zhongzhen Tong, Sifan Sun, Chenghang Li, Jiye Yao, Yulong Qiu, Chao Wang 0094, Zhaohao Wang, Amara Amara, Xiaoyang Lin, Weisheng Zhao 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2024 CIMQ: A Hardware-Efficient Quantization Framework for Computing-In-Memory-Based Neural Network Accelerators
abstract
The novel computing-in-memory (CIM) technology has demonstrated significant potential in enhancing the performance and efficiency of convolutional neural networks (CNNs). However, due to the low precision of memory devices and data interfaces, an additional quantization step is necessary. Conventional NN quantization methods fail to account for the hardware characteristics of CIM, resulting in inferior system performance and efficiency. This article proposes CIMQ, a hardware-efficient quantization framework designed to improve the efficiency of CIM-based NN accelerators. The holistic framework focuses on the fundamental computing elements in CIM hardware: inputs, weights, and outputs (or activations, weights, and partial sums in NNs) with four innovative techniques. First, bit-level sparsity induced activation quantization is introduced to decrease dynamic computation energy. Second, inspired by the unique computation paradigm of CIM, an innovative arraywise quantization granularity is proposed for weight quantization. Third, partial sums are quantized with a reparametrized clipping function to reduce the required resolution of analog-to-digital converters (ADCs). Finally, to improve the accuracy of quantized neural networks (QNNs), the post-training quantization (PTQ) is enhanced with a random quantization dropping strategy. The effectiveness of the proposed framework has been demonstrated through experimental results on various NNs and datasets (CIFAR10, CIFAR100, and ImageNet). In typical cases, the hardware efficiency can be improved up to 222% with a 58.97% improvement in accuracy compared to conventional quantization methods.
Jinyu Bai, Sifan Sun, Weisheng Zhao 0001, Wang Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2024 CIM²PQ: An Arraywise and Hardware-Friendly Mixed Precision Quantization Method for Analog Computing-In-Memory
abstract
Computing-in-memory (CIM) architecture is a promising convolutional neural network (CNN) accelerator known for its highly efficient matrix-vector multiplications (MVMs). However, due to the low-precision computation and limited size of CIM memory arrays, it is necessary to decompose the huge MVMs into smaller subsets. Conventional NN quantization methods overlook the characteristics of CIM hardware, resulting in diminished system performance and efficiency. This paper proposes a mixed precision quantization (MPQ) method based on evolutionary algorithm for CIM-based accelerators, while considering the hardware characteristics of CIM, called CIMPQ, which can automatically generate quantization strategies for NN model to improve the efficiency of CIM systems. Firstly, inspired by the CIM computing paradigm, an array-wise quantization granularity is introduced in the MPQ search space, which can jointly quantize the inputs, weights, and partial sums. Secondly, a production procedure containing fine-grained crossover and progressive adaptive mutation is proposed, which can efficiently explore the search space and speed up the search process. Thirdly, we propose a fast and efficient strategy evaluation method to obtain the performance of quantization strategy on the CIM platform, saving the evaluation time significantly without requiring fine-tuning. Finally, to protect CIM-friendly strategies with lower bit-widths but worse algorithm performance, we propose a strategy selection method based on multi-objective optimization, named qNSGA-III. The effectiveness of the proposed method has been demonstrated through experimental results of various NNs and datasets. For ResNet-18, the hardware efficiency and accuracy can be improved to 117% with 7.05%, 113% with 3.37%, and 119% with 5.78%, on CIFAR-10, CIFAR-100 and ImageNet, respectively, compared to the baseline MPQ method.
Sifan Sun, Jinyu Bai, Zhaoyu Shi, Weisheng Zhao 0001, Wang Kang 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.1
2023 Hierarchical Non-Structured Pruning for Computing-In-Memory Accelerators with Reduced ADC Resolution Requirement
abstract
The crossbar architecture, which is comprised of novel nano-devices, enables high-speed and energy-efficient computing-in-memory (CIM) for neural networks. However, the overhead from analog-to-digital converters (ADCs) substantially degrades the energy efficiency of CIM accelerators. In this paper, we introduce a hierarchical non-structured pruning strategy where value-level and bit-level pruning are performed jointly on neural networks to reduce the resolution of ADCs by using the famous alternating direction method of multipliers (ADMM). To verify the effectiveness, we deployed the proposed method to a variety of state-of-the-art convolutional neural networks on two image classification benchmark datasets: CIFAR10, and ImageNet. The results show that our pruning method can reduce the required resolution of ADCs to 2 or 3 bits with only slight accuracy loss (~0.25 %), and thus can improve the hardware efficiency by 180%.
Wenlu Xue, Jinyu Bai, Sifan Sun, Wang Kang 0001
DATE3