Chunyu Peng

dblp:160/0847 · DBLP profile ↗
← Back
40ranked-venue papers
4as first author
33since 2021 · last 2026
0000-0003-2408-5048ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 36 · 4 first-author · 31 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A 2RW Dual-Port 8T-SRAM Macro with Bitline Leakage Current Tracking and Read-Write Arbitration
Chenghu Dai, Junbo Chen, Zaihang Zhang, Licai Hao, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu
ISCAS7
2026 An offset-compensation capacitor-coupled DRAM sense amplifier with symmetric sensing and high PVT stability
Chenghu Dai, Jing Lv, Yongqi Qin, Chunyu Peng, Xin Li 0099, Yu Liu 0113, Xiulong Wu, Zhi-Ting Lin
Integr.5
2026 MTJ-back-gate SRAM CIM with replica quantization for temperature robust
Yongliang Zhou, Chengxing Dai, Xiulong Wu, Chunyu Peng
Integr.5
2026 Low-temperature-drift voltage reference design using magnetic tunnel junctions
Yongliang Zhou, Yingxue Sun, Wangyong Si, Jingxue Zhong, Weizhe Tan, Chunyu Peng, Xiulong Wu
Integr.7
2026 Analysis and Design of Memory Testing Algorithm for Computing-in-Memory Using MBIST
abstract
Computing-in-memory (CIM), as a novel computing architecture for the future, effectively overcomes the bottlenecks in the von Neumann architecture. The CIM architecture embeds logic into the memory array to reduce the data transfer between the processor and memory. However, embedding logic into the memory array increases the test complexity. In this study, we offer a comprehensive examination of the challenges associated with CIM and introduce a novel March-like test algorithm, named March CC, tailored for CIM chips. Computational elements are added to the read/write operation sequences, combining the tests in memory mode and computing mode into one step, which significantly improves the test efficiency. In comparison to the traditional March C− test algorithm, the proposed March CC test algorithm, with a complexity of only 10 N , enhances the fault coverage from 66.7% to 79.8% for six common single-cell fault (SCF) models and nine common double-cell fault (DCF) models. Furthermore, the March CC algorithm demonstrates good compatibility and is applicable to various memory configurations, such as SRAM, RRAM, and MRAM CIM architectures.
Zhi-Ting Lin, Siyan Li, Qiushi Feng, Changxin Yue, Yuanyang Wang, Yunlong Liu 0006, Yu Liu 0113, Licai Hao, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Xiulong Wu
ACM J. Emerg. Technol. Comput. Syst.11
2026 Time-Domain SRAM-CIM Macro With Dual-Edge Temporal Fused Accumulation for Signed 8-bit Precision MAC
Wenjuan Lu, Xiaobo Gong, Kang Meng, Xiaohang Chen, Jiating Guo, Lijun Guan, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu, Chunyu Peng
IEEE Trans. Circuits Syst. I Regul. Pap.11
2026 A T8T-SRAM Computing-in-Memory Macro for Ternary Deep Neural Networks and Boolean Logic Computations
abstract
Deep neural networks (DNNs) play important roles in artificial intelligence applications and show hungry computility and power demands. Compared with binary neural networks (BNNs), ternary neural networks (TNNs) have higher representation and adaptive abilities and balance the inference accuracy and computing efficiency between DNNs and BNNs. This article proposed a T8T-SRAM computing-in-memory (CIM) macro to achieve Boolean logic operations and MAC operation of ternary activation and ternary weight. The proposed T8T-SRAM bitcell has a separate read and write path, and can avoid the read disturb issue. In Boolean logic operation mode, the T8T-SRAM macro can achievenand,nor,xnor, andxoroperations with redundant rows, reducing the additional reference voltage generation circuit. In the MAC mode, the result is quantized by an embedded column analog-to-digital converter (ADC), which uses activation refresh to reduce weight changing. In 28-nm CMOS technology, under 0.5-V array supply voltage and 0.9-V peripheral supply voltage, simulation results manifest that the MAC results have good linearity, and feasibility of Boolean logic operation. The proposed T8T-SRAM macro realizes MAC operation of 16 ternary activations and 16 ternary weights with 333.99–816.1-TOPS/W energy efficiency and 61.9-TOPS/mm2area efficiency. Using an ResNet-18 network for the inference of MNIST, and CIFAR-10 datasets, the accuracies were 99.06% and 85.76% with a ternary activation and ternary weight.
Chenghu Dai, Zihua Ren, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.7
2026 A Digital FP CIM Macro With Cascaded Row-Elimination for Exponent Comparison and In-Memory Mantissa Sparsity Detection
abstract
Floating-point (FP) computing-in-memory (CIM) addresses the energy efficiency bottleneck of von Neumann architectures and fixed-point CIM in high-accuracy neural network training/inference. However, existing FP CIM designs still suffer from limited exponent-path parallelism and low mantissa-path energy efficiency. This article proposes a novel FP CIM architecture using: 1) a cascaded row-elimination comparison mechanism for single-cycle, high-parallelism exponent max-value comparison; 2) an in-memory sparse feature detection method that skips redundant mantissa computing modules based on different input/weight mantissa patterns to reduce energy consumption; and 3) separated exponent/mantissa CIM modules enabling pipelining and a mantissa bit-extension strategy optimizing accuracy, area, and energy efficiency. Simulation results of a 28-nm 19-kb macro show a 496-MHz operating frequency at 0.9 V and a peak energy efficiency of 12.88 TFLOPS/W.
Wenjuan Lu, Kang Meng, Xiaobo Gong, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.7
2026 A 28-nm 9-kb SRAM Computing-in-Memory Macro With Segmented Charge Sharing for Multimode MAC Operations
abstract
Computing-in-memory (CIM) is a novel approach to solve the von Neumann bottleneck and improve energy efficiency and throughput. This article presents an SRAM-CIM macro based on segmented charge sharing to support multimode multiply-and-accumulate (MAC) operations, including binary weight network (BWN) MAC, ternary weight network (TWN) MAC, and multibit MAC operations. Signed 9-bit input MAC operations can be supported in BWN and TWN networks through multiple cycles. The multibit MAC operations are realized by the cooperation of two cell structures, which eliminates the need to convert negative numbers into complements and reduces the area and power consumption of the chip. Additionally, the proposed segmented charge-sharing scheme increases the speed of accumulation and improves the overall energy efficiency of the chip. The proposed 9-kb macro is implemented in 28-nm CMOS technology with an energy efficiency of 30.74 TOPS/W and an area efficiency of 1.15 TOPS/mm2in multibit MAC mode. At the system level, a ResNet-based model achieves an accuracy of 94.59% on the CIFAR-10 dataset and 74.64% on the CIFAR-100 dataset, demonstrating the effectiveness of the proposed CIM architecture for practical neural network inference. Compared to prior designs, it not only supports additional computing modes but also demonstrates notable improvements in both energy and area efficiency.
Wenjuan Lu, Lening Tan, Tianchen Xue, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.5
2026 A 28-nm 16-Kb SRAM Computing-in-Memory Macro With Dual Bitline Computing for Boolean Logic Operation and BWN MAC Operation
abstract
By integrating computation directly within memory structures, computing-in-memory (CIM) effectively mitigates the von Neumann bottleneck, offering significant improvements in energy efficiency and system throughput. This letter presents a 16-kb static random access memory (SRAM) CIM macro in 28-nm complementary metal–oxide–semiconductor (CMOS), featuring a novel dual bitline computing cell (DBCC) architecture. The design supports three operational modes: conventional memory access, Boolean logic operations, and binary weight network (BWN) multiply-and-accumulate (MAC) computations. The DBCC architecture addresses key critical limitations of prior works by: 1) storing both positive and negative weights within a single 6T-SRAM cell, improving array utilization by$2\times $compared to segregated storage schemes; 2) achieving consistent discharge rates for both polarities through symmetric pull-down paths, enhancing calculation accuracy [INL <3 least significant bit (LSB) across corners] and process variation tolerance; and 3) enabling differential voltage-based Boolean logic operations without external reference voltage generation, thereby reducing peripheral circuitry. The design features a compact 2:1 multiplexer (2:1 MUX)-based pulsewidth modulator (PWM) for 5-b signed input encoding and demonstrates 55.8–204.8-TOPS/W energy efficiency in 28-nm CMOS. Experimental results on a 16-kb array show stable operation across PVT variations while supporting both memory functions and in situ computation for BWNs. System-level evaluation demonstrates its practicality, achieving competitive inference accuracies of 97.59% on MNIST and 89.08% on CIFAR-10.
Wenjuan Lu, Lubin Xiang, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.5
2026 A Floating-Point SRAM Computing-in-Memory Macro Using Digital-Domain Structure for CNNs
Wenjuan Lu, Xiaobo Gong, Xiulong Wu, Chunyu Peng
IEEE Trans. Very Large Scale Integr. Syst.8
2026 A Capacitor Discharge-Based SRAM CIM Macro Based on Hybrid-Domain for Convolutional Neural Networks
abstract
Compute-in-memory (CIM) is increasingly recognized as an effective hardware accelerator for convolutional neural networks (CNNs). This work proposes a hybrid-domain CIM design using: 1) a multibit compute unit (MBCU) structure that realizes the multiplication operation of 2-bit input and 4-bit weight through the transistor-size-weighted capacitor discharge on the bitline; 2) a hybrid-domain quantization scheme (HDQS) of “time-domain + voltage-domain,” which integrates the high energy efficiency of time-domain quantization with the low-delay advantages of the voltage-domain quantization, and enhances the quantization accuracy through the combined effect of the process tracking module and the reference signal module; 3) the CIM circuit design, layout drawing and simulation verification of hybrid-domain static random access memory (SRAM) were realized by 28-nm CMOS technology, results show that the circuit supports 8-bit multiply–accumulate (MAC) operation, and full-precision quantization in the hybrid-domain form can achieve the optimal energy efficiency of 249.7 TOPS/W per bit at 0.7 V, and area efficiency of 4.29 TOPS/mm2per bit. Furthermore, the integration of the circuits with the VGG-16 network has been demonstrated to yield an inference accuracy of 90.52% in the CIFAR-10 dataset.
Bin Qiang, Yongliang Zhou, Xiulong Wu, Chunyu Peng
IEEE Trans. Very Large Scale Integr. Syst.5
2026 A CIM Macro Embedded With Sign Operations for Parallel Signed Multibit Multiplication-and-Accumulation Using Hybrid Cell Array
Jin Zhang 0036, Zhongzhen Tong, Qiang Zhao 0007, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.4
2026 Enhanced monocular depth estimation via semantic fusion and planar constraints
Chunyu Peng, Zhensong Li, Shoubiao Tan
Vis. Comput.2
2025 MTJ based temperature compensated beta multiplier Voltage Reference
abstract
This article mainly explores the interaction between CMOS devices and Magnetic Tunnel Junction (MTJ) devices in terms of temperature characteristics, aiming to achieve a CMOS beta multiplier circuit that combines low power consumption and wide temperature adaptability, making it a stable reference voltage source. The proposed design utilizes the Tunneling Magneto Resistance (TMR) effect of MTJ to compensate for the performance mismatch caused by temperature changes in CMOS. This design adopts TSMC 28nm CMOS craft, which can generate a reference voltage with a linearity of 0.57%/V, a temperature coefficient of 43.6ppm/°C, and a stable voltage of 441.6mV at a minimum supply voltage of 0.6V and a temperature range of 10~110 °C. The noise of this voltage at a frequency of 10Hz is 69.7uV/sqrt (Hz), the power suppression ratio is -62.1dB, and the power consumption is 4.092nW.
Yongliang Zhou, Yingxue Sun, Jingxue Zhong, Chengxing Dai, Weizhe Tan, Chunyu Peng, Xin Li 0099, Zhi-Ting Lin, Xiulong Wu
ISCAS6
2025 MTJ based Temperature-Adaptive VCO (TAVCO) for Compensating CP-PLL Frequency Drift
abstract
The Charge Pump Phase-Locked Loop (CP-PLL) is a commonly utilized component in contemporary mixed-signal electronic systems. It is widely employed for clock generation, synchronization, and frequency synthesis in both digital and wireless functionalities. However, the frequency accuracy of oscillators can be adversely affected by variations in frequency across a broad temperature range. To address this issue, the voltage-controlled oscillator designed in this study employs a four-stage differential delay structure, chosen for its simple circuit architecture, favorable control linearity, and low noise characteristics. This research integrates the temperature behaviors of Complementary Metal-Oxide-Semiconductor (CMOS) and Magnetic Tunnel Junction (MTJ) technologies, utilizing 28nm CMOS technology to enhance the frequency stability of ring oscillators effectively. Simulation results indicate that frequency drift is reduced by 92% within the temperature range of -80°C to 125°C.
Yongliang Zhou, Jingxue Zhong, Yingxue Sun, Chengxing Dai, Weizhe Tan, Chunyu Peng, Wenjuan Lu, Xin Li 0099, Zhi-Ting Lin, Xiulong Wu
ISCAS6
2025 A Low-Cost and Triple-Node-Upset Self-Recoverable Latch Design With Low Soft Error Rate
abstract
With the decrease in feature size of transistors, latches are more sensitive to single-event multiple node upset (MNU), including double node upset (DNU) and triple node upset (TNU). However, the reported TNU self-recoverable (TNUR) latches are facing problems with large areas and power consumption. Based on the polarity design, this article proposes a low-cost TNUR latch (LCTRL) with a low soft error rate (SER) in 28-nm CMOS technology. The proposed LCTRL mainly consists of four interlocked modules and a clock-gated inverter. Compared with the state-of-the-art TNUR latches, including LCTNURL, IHTRL, FATNU, and TRLW, the power consumption, D-Q delay, CLK-to-Q delay, area, and the power-delay–area product (PDAP) of the proposed LCTRL are reduced by 55.09%, 38.64%, 42.93%, 44.65%, and 83.50%, respectively. Due to the polarity design, the SER of the proposed LCTRL is the smallest among compared latches, which suggests that the proposed LCTRL is suitable for use in radiation environments.
Licai Hao, Lang Tian, Hao Wang 0239, Shiyu Zhao 0004, Qiang Zhao 0007, Chunyu Peng, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.6
2025 Full-Array Boolean Logic CIM Macro With Self-Recycling 10T-SRAM Cell for AES Systems
abstract
Computing in memory (CIM), which alleviates the need to transfer a large amount of data between processor and memory, significantly reducing latency and energy consumption, is a promising new computing architecture for addressing the von Neumann bottleneck problem. This article proposes a CIM array structure composed of self-recycling 10T static random access memory (SRAM) cells, which can realize orthogonal data writing, and multiple Boolean logical operations for the entire array. The self-recycling and full-array activation characteristics are extremely suitable for accelerating diverse data processing algorithms such as the Advanced Encryption Standard (AES). A 4-kb SRAM is implemented in 55-nm CMOS technology to verify the effectiveness of the design. Compared with other state-of-the-art architectures, the throughput and the operating frequency of the proposed CIM macro are increased to 843 GOPS/kb ($2.64\times $) and 823.7 MHz ($2.6\times $), respectively. The energy efficiency reaches 246.9 TOPS/W. When applied to the AES, the energy consumption is 35.77% less than the digital CIM architecture that is not self-recycling.
Xin Li 0099, Lintao Chen, Yang Lou, Baofa Wu, Jiajun Long, Yongliang Zhou, Chunyu Peng, Xiulong Wu, Zhi-Ting Lin
IEEE Trans. Very Large Scale Integr. Syst.9
2025 A 28-nm 9T1C SRAM-Based CIM Macro With Hierarchical Capacitance Weighting and Two-Step Capacitive Comparison ADCs for CNNs
abstract
In the realm of charge-domain computing-in-memory (CIM) macros, reducing the area of capacitor ladder and analog-to-digital converter (ADC) while maintaining high throughput remains a significant challenge. This brief introduces an adjustable-weight CIM macro designed to enhance both energy efficiency and area efficiency for convolutional neural networks (CNNs). The proposed architecture uses: 1) a customized 9T1C bit cell for sensing margin improvement and bidirectional decoupled read ports; 2) a hierarchical capacitance weighting (HCW) structure that achieves a weight accumulation of 1/2/4 bits with less capacitance area and weighting time; and 3) a two-step capacitive comparison ADCs (TC-ADCs) readout scheme to improve area efficiency and throughput. The proposed 8-kb static random address memory (SRAM) CIM macro is implemented using 28-nm CMOS technology. It can achieve an energy efficiency of 224.4 TOPS/W and an area efficiency of 21.894 TOPS/mm2, and the accuracies on MNIST, CIFAR-10, and CIFAR-100 datasets are 99.67%, 89.13%, and 67.58% with a 4-b input and 4-b weight.
Zhi-Ting Lin, Runru Yu, Miao Long, Yu Liu 0113, Jianxing Zhou, Qingchuan Zhu, Yue Zhao 0029, Lintao Chen, Chunyu Peng, Qiang Zhao 0007, Xin Li 0099, Chenghu Dai, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.11
2025 High-Reliability and High-Throughput CIM 10T-SRAM for Multiplication and Accumulation Operations With 274.3 GOPS and 200-237.5 TOPS/W
abstract
Artificial intelligence (AI) is extensively applied in natural language processing, image matching, and image recognition, with convolutional neural networks (CNNs) being crucial. Computing-in-memory (CIM) utilizing static random access memory (SRAM) can enhance the CNN performance. However, this faces issues such as multibit signed data processing, read corruption of traditional SRAM arrays, and increased area overhead due to increased capacitor weighting. This article proposes a 10T-SRAM macro tailored for CNN multiply-accumulate calculation (MAC) computation in image processing. It enables high-throughput full-array operations, with added dual ports facilitating input of multibit data with signed bits. The 10T-SRAM cell features a read-write separation channel, mitigating read disturbance issues seen in dual-port 8T-SRAM arrays or 6T-SRAM arrays. Incorporating redundant columns in the array for charge sharing and weighting conserves area and boosts circuit reliability. In the 28-nm CMOS simulation environment, the proposed architecture achieves a throughput of 274.3 GOPS and an energy efficiency of 200–237.5 TOPS/W, surpassing literature-reported figures by several times.
Wenjuan Lu, Lubin Xiang, Chunyu Peng, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.4
2025 A 28 nm Dual-Mode SRAM-CIM Macro With Local Computing Cell for CNNs and Grayscale Edge Detection
abstract
With the rise of artificial intelligence (AI), neural network applications are growing in demand for efficient data transmission. The traditional von Neumann architecture can no longer keep pace with modern technological needs. Computing-in-memory (CIM) is proposed as a promising solution to address this bottleneck. This work introduces a local computing cell (LCC) scheme based on compact 6T-SRAM cells. The proposed circuit aims to enhance energy efficiency and reduce power consumption by reusing the LCC. The LCC circuit can perform the multiplication of a 2-bit input with a 1-bit weight, which can be applied to convolutional neural networks (CNNs) with the multiply-accumulate (MAC) operations. Through circuit reuse, it can also be used for multibit multiply operations, performing 2-bit input multiplication and 1-bit weight addition, which can be applied to grayscale edge detection in images. The energy efficiency of the SRAM-CIM macro achieves an energy efficiency of 46.3 TOPS/W under MAC operations with input precision of 8-bits and weight precision of 8-bits, and up to 389.1–529.1 TOPS/W under the calculation in one subarray with an input precision of 2-bits and a weight precision of 1-bit. The estimated inference accuracy on CIFAR-10 datasets is 90.21%.
Chunyu Peng, Xiaohang Chen, Mengya Gao, Jiating Guo, Lijun Guan, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.1
2025 A 28-nm Cascode Current Mirror-Based Inconsistency-Free Charging-and-Discharging SRAM-CIM Macro for High-Efficient Convolutional Neural Networks
abstract
Computing-in-memory (CIM) is an emerging approach to alleviate the von Neumann bottleneck and enhance energy efficiency and throughput. This brief introduces a 16-Kb static random access memory (SRAM) CIM macro for convolutional neural networks (CNNs), featuring a cascode current mirror-based inconsistency-free computing circuits (CICCs). The bias voltage of CICC is provided by a cascode current mirror (CCM) circuit. The proposed architecture improves the consistency and linearity of bitline (BL) charge and discharge rates in the analog current domain, enhancing computational accuracy. Additionally, the charge and discharge on the BLs represent the positive or negative calculation result, eliminating the need for extra encoding and logic circuits to handle sign bits. The SRAM-CIM macro achieves an energy efficiency of 59.1–134.0 TOPS/W and a throughput of 0.41 TOPS in a 28-nm CMOS technology, and the estimated inference accuracy on MNIST and CIFAR-10 datasets is 96.5% and 91.4%, respectively, with 5-bit input precision and 1-bit weight precision.
Chunyu Peng, Jiating Guo, Shengyuan Yan, Xiaohang Chen, Wenjuan Lu, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.1
2025 A High-Performance and High-Robustness Triple-Node-Upset Tolerant Latch Based on Redundant-Node Hardening
abstract
In response to the issues of high cost, large overhead, and limited node fault tolerance in current latch hardening techniques, this article proposes a latch circuit resistant to triple-node-upset (TNU) based on redundant-node hardening technology. This latch comprises eight 1P2N modules interlocked, with its output isolated by two levels of C-elements (CEs), achieving tolerance to TNU. The performance of the redundant-node reinforcement TNU tolerant latch (RNRTTL) was simulated and verified using CMOS 65 nm technology. The simulation results indicate that the RNRTTL circuit has a D-Q delay of 14.14 ps, static power consumption of$4.03~\mu $w, an area of$32.87~\mu $m2, and an area-static power-D–Q delay-product (APDP) of 1873, respectively. Compared to the triple-node upset tolerant latches TTLL, TNU-latch, TNURL, and HLTNURL reported in the current literature, the proposed latch demonstrates an average reduction of 219.9%, 164.9%, 150.7%, and 2464.8% in D-Q delay, static power consumption, area, and APDP, respectively, indicating that the RNRTTL latch has superior comprehensive performance; furthermore, a series of 2000 Monte Carlo (MC) simulations on the node group$\langle $Q, X0, X$8\rangle $reveal that the proposed latch circuit possesses good stability, making it suitable for harsh radiation environments.
Qiang Zhao 0007, Qingyi Liu, Licai Hao, Xin Li 0099, Shengyue Zhang, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.7
2024 SRAM-Based Digital CIM Macro for Linear Interpolation and MAC
abstract
Linear interpolation is widely used in algorithms such as image segmentation, but the existing compute-in-memory (CIM) architectures cannot satisfy the needs of linear interpolation. This paper proposes a CIM macro based on static random-access memory (SRAM) that implements linear interpolation for the first time. Swing multiplication and accumulation are proposed in this paper for linear interpolation operations. In addition, the proposed circuit can be used for multiply-and-accumulate (MAC) operation and supports parallel updating and computing. A new adder tree by reused the adders inside the multiplier to implement the accumulation operation. The proposed circuit can improve the area efficiency as no additional adder tree circuit is required. The design is implemented in a 28 nm process and area efficiency can achieve 0.059 TOPS/mm2for MAC operation. When operating with an 8-bit linear interpolation, this CIM macro achieves access times of 6−9 ns and energy efficiencies of 11.13−17.72 TOPS/W. Under MAC operation with 8-bit inputs and weights, this CIM macro achieves access time of 6.36−9.47 ns and energy efficiency of 4.61−8.36 TOPS/W for 19-bit outputs.
Zhi-Ting Lin, Yunlong Liu 0006, Yaling Wang, Yue Zhao 0029, Chunyu Peng, Xiulong Wu
ISCAS5
2024 A Timing-Shared Adaptive Sensing Methodology for Low-Voltage SRAM
abstract
Lowing static random access memory (SRAM) supply voltage could highly improve energy efficiency, yet energy efficiency still not attains optimal point due to the constraint of the weakest bit-cell, especially in low-voltage SRAM. Adaptive sensing methodology is proposed for the challenge and consists of four elements: Switch unit is built to implement cross-sensing operation, Timing-Shared Decoupling Latch Sense Amplifier (TS-DLSA) allows rapid and successive sensing, the judging module is utilized to trigger the FLAG signal which is for adaptive timing controller to cut off word line (WL). Energy efficiency was obtained by compressing the activation delay of WL compared to global timing scheme. The proposed adaptive sensing methodology was performed with TSMC 28-nm CMOS process, and evaluated in SRAM array of 128x128, 256x256, 512x512, and 1024x1024. Monte Carlo simulation results are formed to confirm that, compared to the global timing scheme, the proposed adaptive sensing methodology has reduced WL activation delay by 75.1%∼41.8% and read operation energy overhead by 77.5%∼30.9% from 0.6V to 1.2V. Compared to the current-latched sense amplifier with a footswitch (FS-CLSA) with proposed adaptive sensing methodology, TS-DLSA with proposed scheme have reduced the WL activation by 4.1%∼10.5% and read operation energy overhead by 15.4%∼23.8% from 256x256 to 2048x2048 at 0.6V. The more cells mounted on the BL, the higher energy revenue gains.
Yongliang Zhou, Saiai Wu, Wenjuan Lu, Chunyu Peng, Xin Li 0099, Xiulong Wu
ISCAS7
2024 Timing Optimization Model and PVT Tracked Scheme for STT-MRAM Voltage-Mode Sense
abstract
The impact of process variations on the read operation of low-voltage STT-MRAM becomes severe, posing a challenge in determining the optimal sensing timing of the sense amplifier. This study investigates techniques for refining the timing scheme of sensing circuits in order to improve the sensing reliability of the STT-MRAM. The supply voltage$V_{DD}$, the Tunneling Magnetoresistance Ratio TMR, the low resistance state of bit-cell$R_{P}$, and the parasitic capacitance of bit-line$C_{BL}$are analyzed along with the voltage sense amplifier (VSA) involved in sensing yield. We develop a timing model through theoretical analysis to determine the optimal VSA enable signal (SAE). In addition, an innovative Process-Voltage-Temperature (PVT) tracking scheme is proposed that can track the optimal VSA enable signal (SAE) and suppress timing variations. Monte-Carol simulation in the 28nm CMOS and magnetic tunnel junction (MTJ) process confirms that the combined scheme significantly enhances the robustness of sensing operation. The proposed scheme improves yield by 20% to 35%, reduces power consumption by 43% to 63%, and reduces read access delay by 47% to 59% compared to conventional sensing schemes at 0.6V supply voltage.
Yongliang Zhou, Yingxue Sun, Chengxing Dai, Jingxue Zhong, Xiulong Wu, Chunyu Peng
IEEE Trans. Circuits Syst. I Regul. Pap.10
2024 A CFMB STT-MRAM-Based Computing-in-Memory Proposal With Cascade Computing Unit for Edge AI Devices
abstract
The application of non-volatile memory technology is increasingly attractive for Computing-in-memory (CIM) owing to high integration density and negligible standby power consumption. This study proposes an spin-transfer-torque (STT) magnetic random access memory (MRAM) based CIM macro which incorporates following innovative features: 1) cross-feedback margin-boost (CFMB) scheme to enable robust and fast reading operations against process variation and limited Tunneling Magnetoresistance Ratio (TMR); 2) cascade computing units (CCU) and related design method for efficient and stable multi-bit multiply-and-accumulate (MAC) operation; and 3) dual computing mode scheme and resolution adjustable quantization module to optimize energy efficiency and operating speed. The post-simulations are performed under 28nm CMOS&MTJ technology. The results demonstrate the achievement in energy efficiency of 36.4 TOPS/W while performing MAC operations with up to 16-bit weights, 4-bit inputs, and 22-bit outputs.
Yongliang Zhou, Chenghu Dai, Licai Hao, Chunyu Peng, Hao Cai 0001, Xiulong Wu
IEEE Trans. Circuits Syst. I Regul. Pap.8
2024 Low-Cost and Highly Robust Quadruple Node Upset Tolerant Latch Design
abstract
This article proposes an exceptionally reliable and low-cost quadruple node upset tolerant latch ($LC$-QNUTL) suitable for the 65 nm CMOS technology. The innovative$LC$-QNUTL latch is primarily composed of three soft-error-immune (SEI) static random-access memory (SRAM) cells and a triple-level C-element (CE) unit, which includes five two-input CE and a clock-gating (CG)-based two-input CE. The SEI SRAM cell utilizes polarity hardening technology and source-isolation technology, significantly reducing the number of sensitive nodes and enhancing the latch’s stability. By using the high-speed transmission gate (TG) technology and stacked structures, the proposed latch offers minimal overhead in terms of delay and power consumption, yielding an improved power delay area product (PDAP). When compared to contemporary quadruple node upset (QNU)-tolerant latch designs (including HLMR, 4NUHL, and LDAVPM), the new design offers substantial improvements—29.53% less delay, 80.09% reduced power consumption, 58.52% smaller silicon area, and 433.43% improved comprehensive PDAP on average. Furthermore, simulation results demonstrate that the$LC$-QNUTL latch exhibits reduced sensitivity to process, voltage, and temperature (PVT) variations, thus providing superior reliability, which makes it an ideal choice for safety-critical applications.
Licai Hao, Yaling Wang, Yunlong Liu 0006, Shiyu Zhao 0004, Wenjuan Lu, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.8
2024 Soft-Error-Immune Quadruple-Node-Upset Tolerant Latch Based on Polarity Design and Source-Isolation Technologies
abstract
A soft-error-immune quadruple-node-upset tolerant latch (SEI-QNUTL) with a low delay and high performance is proposed using 65-nm CMOS technology. The proposed SEI-QNUTL design consists of three soft-error-immune static random access memory (SEI-SRAM) cells. Furthermore, each SEI-SRAM cell employs polarity design and source-isolation technology to reduce the number of sensitive nodes and enhance the reliability of the latch. Compared with state-of-the-art quadruple-node-upset (QNU) tolerant latches [including high-performance and low-cost single-event multiple-node-upsets resilient (HLMR), QNU tolerant latch (QNUTL), and Latch Design and Algorithm-based Verification Protected against Multiple-Node-Upsets (LDAVPM)], the proposed SEI-QNUTL design reduces (on average) the area, delay, and area-power-delay-product (APDP) by 47.0%, 25.0%, 46.5%, and 66.3%, respectively. Extensive variation analysis validates that the SEI-QNUTL design is less sensitive to process, voltage, and temperature (PVT) variations regarding power consumption and delay. Furthermore, Monte Carlo (MC) simulations show that the proposed latch exhibits high reliability when performing data storage. Compared with the existing latches, the SEI-QNUTL design makes a good tradeoff among delay, power, and area, and it can thus be used in safety-critical applications.
Licai Hao, Chenghu Dai, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng, Yongliang Zhou, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.6
2023 MS3DAAM: Multi-scale 3-D Analytic Attention Module for Convolutional Neural Networks
Yincong Wang, Shoubiao Tan, Chunyu Peng
ICONIP (1)3
2023 High Restore Yield NVSRAM Structures With Dual Complementary RRAM Devices for High-Speed Applications
abstract
Static random access memory (SRAM) plays a key role in the overall performance of electronic systems because of its rapid data processing and transmission speed; however, when the system power supply is cut off, the data stored in the nodes are lost. Thus, this article proposes four nonvolatile SRAM (NVSRAM) cells that use unilateral or bilateral structures with dual complementary series resistive random access memory (RRAM) devices. It is found that the read, write, and hold static noise margins (HSNMs) are comparable with those of the standard 6T-SRAM. Moreover, the store and restore operations operate in parallel at high speed. The store operation delay is only 6 ns for unilateral structures and 5 ns for bilateral structures, and the restore delay is only 10 ns for unilateral structures and 6 ns for bilateral structures. The maximum power consumption among the four structures for storing and restoring a “1” are 1.545 pJ/bit and 134.5 fJ/bit, respectively. Furthermore, the dual complementary series resistor structures can achieve a high restore yield at a resistance ratio of 1.5. Therefore, a high restore yield can be achieved even with large resistance fluctuations caused by the voltage, time, and process.
Zhi-Ting Lin, Xiulong Wu, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng
IEEE Trans. Very Large Scale Integr. Syst.7
2023 A Fully Digital SRAM-Based Four-Layer In-Memory Computing Unit Achieving Multiplication Operations and Results Store
abstract
The separation of memory and arithmetic logic unit (ALU) in the von Neumann computing architecture hinders the development of big data and high-performance computing. In-memory computing (IMC) as a new computation method significantly reduces the latency and power consumption of data processing. In this study, we propose a fully digital static random access memory (SRAM)-based IMC architecture, which has the following advantages: 1) it simplifies multiplication to multicycle addition operations, reuses logic cells, and reduces hardware overhead; 2) by adding a pair of nMOS transistors to achieve internal write-back, the computational efficiency is improved, and at the same time, the final result of the multiplication can be stored locally, eliminating the need to read the computational result immediately; and 3) this scheme can be easily expanded to multiplication operations with different bit widths, which provides good scalability. A 4-kb SRAM-IMC macro chip is manufactured using the SMIC 55-nm technology to realize 4-bit multiplication, with an energy efficiency of 51.4 TOPS/W (0.9 V) and a throughput of 234.3 GOPS/mm2. The proposed multiplication–accumulation architecture is applied to a neural network, which achieves 98.7% accuracy with the Mixed National Institute of Standards and Technology database (MNIST) dataset.
Zhi-Ting Lin, Shaoying Zhang, Jianping Xia, Yunwei Liu, Kefeng Yu, Zhongzhen Tong, Xiulong Wu, Wenjuan Lu, Chunyu Peng, Qiang Zhao 0007
IEEE Trans. Very Large Scale Integr. Syst.14
2022 Configurable Memory With a Multilevel Shared Structure Enabling In-Memory Computing
abstract
Frequent to-and-from data transfers in the von Neumann architecture limit the overall throughput. One of the promising approaches used to overcome von Neumann bottleneck is in-memory computing (IMC) that aims to embed computing in memory to reduce the transfer of memory-processor data. This study proposes a configurable 6-transistor (6T) static random access memory (SRAM) array with a multilevel shared structure for IMC. A multilevel shared structure can effectively improve the utilization rate of the module. In addition to the conventional SRAM operation, the configurable structure can also perform the sum of absolute differences (SAD) and Hamming distance (HD) calculations. To quickly identify the minimum value among multiple calculation results, a four-input sense amplifier (SA) is proposed. The performance of the proposed memory is simulated in a 65-nm CMOS process. The post-layout simulation results show good linearity of the multirow read in the SAD and HD modes. The mean time required by the four-input SA to obtain the result is 190 ps. The SAD and HD calculations yield consumptions of 67.44 fJ/byte and 0.64 fJ/bit, respectively, at 0.8 V. Furthermore, a single column-sharing comparator consumes 2.78 and 3.41 pJ at 0.8 V in the SAD and HD modes, respectively.
Yue Zhao 0029, Zhi-Ting Lin, Xiulong Wu, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng, Zhongzhen Tong, Junning Chen
IEEE Trans. Very Large Scale Integr. Syst.6
2020 Multiple Sharing 7T1R Nonvolatile SRAM With an Improved Read/Write Margin and Reliable Restore Yield
abstract
This article presents a new resistive random access memory (RRAM)-based average 7T1R nonvolatile static random access memory (nvSRAM). This multiple sharing (MS) 7T1R uses MS schemes in which some of the transistors play various roles. Therefore, the MS-7T1R can perform a decoupled read, enhance write capability, and improve the restore yield with a small area cost. Furthermore, the MS-7T1R can offer two alternative modes, i.e., a high speed and a stable mode. Compared with existing technologies, such as the previous 6T-based nvSRAMs, the results show that the proposed architecture provides a remarkable restore yield and ~154% improvement in the read static noise margin (at TT corner and stable mode). In addition, the read delay improves by ~23% (at TT corner and high-speed mode). The write “1” problem of the single bitline is effectively resolved with our proposed write strategy. The static write margin of “1” is improved by ~88.6% compared with the conventional 6T (β = 4) at a power of 1.2 V. In addition, dynamic power is effectively reduced by the use of the single bitline and sub-word-line driver technology.
Zhi-Ting Lin, Yong Wang 0038, Chunyu Peng, Xiulong Wu, Junning Chen
IEEE Trans. Very Large Scale Integr. Syst.3
2020 In-Memory Computing With Double Word Lines and Three Read Ports for Four Operands
abstract
The von Neumann architecture is approaching its limits in terms of scalability and power consumption. In-memory computation is a possible approach to mitigate this limitation. This brief proposes a configurable 8T static random access memory (SRAM) cell with double word lines and three read ports for in-memory computing. In addition to the normal SRAM function, XOR/XNOR and compound Boolean logic operations of three or four operands, such as AND-OR, AND-OR-INVERT, OR-AND, and OR-AND-INVERT, can be performed in one cycle by fully utilizing the three read ports to obtain 13.2-fJ/bit consumption at 0.6 V. The logic operation frequency is 793 MHz at 1.2 V. The proposed SRAM effectively resolves the bottleneck of the existing in-memory computation schemes that only support compound Boolean logic operations with more than two cycles. In addition, the proposed SRAM array scheme can be configured and used as a binary content-addressable memory or a ternary content-addressable memory for searching operations; it achieves 0.24 fJ/search/bit at 0.6 V in the worst case. At 1.2 V, the searching frequency is up to 813 MHz when searching 128 bits with 65-nm technology.
Zhi-Ting Lin, Honglan Zhan, Chunyu Peng, Wenjuan Lu, Xiulong Wu, Junning Chen
IEEE Trans. Very Large Scale Integr. Syst.4
2020 Novel Write-Enhanced and Highly Reliable RHPD-12T SRAM Cells for Space Applications
abstract
In this brief, we proposed, based on the polarity upset mechanism of single-event transient voltage of n-channel metal-oxide-semiconductor (nMOS) transistors, a novel radiation hardened by polar design (RHPD) 12T SRAM cell to enhance the reliability and operation speed for space applications. Simulation results in Semiconductor Manufacturing International Corporation (SMIC) 65-nm CMOS commercial standard process show that the proposed RHPD-12T cell can tolerate all single-node upsets. Meanwhile, compared with We-QUATRO, QUATRO, and dual interlocked storage cell (DICE), the write speed of the proposed cell can be reduced by ~41.8 and ~35.3%, and the static power consumption is reduced by ~41.6 and ~46.3%, respectively. Monte Carlo (MC) simulation has proved that under high frequency and low supply (0.6 V) voltage, RHPD-12T has the minimum write failure probability compared with five other SRAM cells.
Qiang Zhao 0007, Chunyu Peng, Junning Chen, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.2
2019 Radiation-Hardened 14T SRAM Bitcell With Speed and Power Optimized for Space Application
abstract
In this paper, a novel radiation-hardened 14-transistor SRAM bitcell with speed and power optimized [radiation-hardened with speed and power optimized (RSP)-14T] for space application is proposed. By circuit- and layout-level optimization design in a 65-nm CMOS technology, the 3-D TCAD mixed-mode simulation results show that the novel structure is provided with increased resilience to single-event upset as well as single-event-multiple-node upsets due to the charge sharing among OFF-transistors. Moreover, the HSPICE simulation results show that the write speed and power consumption of the proposed RSP-14T are improved by ~65% and ~50%, respectively, compared with those of the radiation hardened design (RHD)-12T memory cell.
Chunyu Peng, Jiati Huang, Changyong Liu, Qiang Zhao 0007, Songsong Xiao, Xiulong Wu, Zhi-Ting Lin, Junning Chen, Xuan Zeng 0001
IEEE Trans. Very Large Scale Integr. Syst.1
2018 Average 7T1R Nonvolatile SRAM With R/W Margin Enhanced for Low-Power Application
abstract
A new average 7T1R nonvolatile SRAM for low-power application is presented in this brief, which improves the read and write margin (RM/WM), as well as the restore energy, simply by using the source switch transistor. Simulation results demonstrate that the RM and WM will be improved by ~23% and ~73%, respectively, and the energy consumption will be decreased by ~63% for low-resistance state restoration, compared with the prior art initialization-and-overwrite-7T1R at nMOS typical corner and pMOS typical corner in Taiwan Semiconductor Manufacturing Company's 65-nm technology. In addition, with the column-shared structure, the area penalty is cheerfully acceptable.
Chunyu Peng, Songsong Xiao, Wenjuan Lu, Xiulong Wu, Junning Chen, Zhi-Ting Lin
IEEE Trans. Very Large Scale Integr. Syst.1
2015 Image-to-class distance ratio: A feature filtering metric for image classification
Shoubiao Tan, Li Liu 0004, Chunyu Peng, Ling Shao 0001
Neurocomputing3
2015 Multi-stage dual replica bit-line delay technique for process-variation-robust timing of low voltage SRAM sense amplifier
abstract
A multi-stage dual replica bit-line delay (MDRBD) technique is proposed for reducing access time by suppressing the sense-amplifier enable (SAE) timing variation of low voltage static random-access memory (SRAM) applications. Compared with the traditional technique, this strategy, using statistical theory, reduces the timing variation by using multi-stage ideas, meanwhile doubling the replica bit-line (RBL) capacitance and discharge path simultaneously in each stage. At a supply voltage of 0.6 V, the simulation results show that the standard deviations of the SAE timing and cycle time with the proposed technique are 69.2% and 47.2%, respectively, smaller than that with a conventional RBL delay technique in TSMC 65 nm CMOS technology (Taiwan Semiconductor Manufacturing Company, Taiwan).
Shoubiao Tan, Wenjuan Lu, Chunyu Peng, Zhengping Li, Youwu Tao, Junning Chen
Frontiers Inf. Technol. Electron. Eng.3