Xiulong Wu

dblp:28/2200 · DBLP profile ↗
← Back
48ranked-venue papers
0as first author
39since 2021 · last 2026
0000-0002-5012-2570ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 46 · 39 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 A 2RW Dual-Port 8T-SRAM Macro with Bitline Leakage Current Tracking and Read-Write Arbitration
Chenghu Dai, Junbo Chen, Zaihang Zhang, Licai Hao, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu
ISCAS10
2026 A Floating-Point CIM Macro Featuring Asymmetric Exponent Encoding and Adaptive Mantissa Truncation for High-Efficiency AI Edge Computing
Zhi-Ting Lin, Rongtao Li, Yu Liu 0113, Xin Li 0099, Xiulong Wu
ISCAS8
2026 An offset-compensation capacitor-coupled DRAM sense amplifier with symmetric sensing and high PVT stability
Chenghu Dai, Jing Lv, Yongqi Qin, Chunyu Peng, Xin Li 0099, Yu Liu 0113, Xiulong Wu, Zhi-Ting Lin
Integr.8
2026 MTJ-back-gate SRAM CIM with replica quantization for temperature robust
Yongliang Zhou, Chengxing Dai, Xiulong Wu, Chunyu Peng
Integr.4
2026 Low-temperature-drift voltage reference design using magnetic tunnel junctions
Yongliang Zhou, Yingxue Sun, Wangyong Si, Jingxue Zhong, Weizhe Tan, Chunyu Peng, Xiulong Wu
Integr.8
2026 Analysis and Design of Memory Testing Algorithm for Computing-in-Memory Using MBIST
abstract
Computing-in-memory (CIM), as a novel computing architecture for the future, effectively overcomes the bottlenecks in the von Neumann architecture. The CIM architecture embeds logic into the memory array to reduce the data transfer between the processor and memory. However, embedding logic into the memory array increases the test complexity. In this study, we offer a comprehensive examination of the challenges associated with CIM and introduce a novel March-like test algorithm, named March CC, tailored for CIM chips. Computational elements are added to the read/write operation sequences, combining the tests in memory mode and computing mode into one step, which significantly improves the test efficiency. In comparison to the traditional March C− test algorithm, the proposed March CC test algorithm, with a complexity of only 10 N , enhances the fault coverage from 66.7% to 79.8% for six common single-cell fault (SCF) models and nine common double-cell fault (DCF) models. Furthermore, the March CC algorithm demonstrates good compatibility and is applicable to various memory configurations, such as SRAM, RRAM, and MRAM CIM architectures.
Zhi-Ting Lin, Siyan Li, Qiushi Feng, Changxin Yue, Yuanyang Wang, Yunlong Liu 0006, Yu Liu 0113, Licai Hao, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Xiulong Wu
ACM J. Emerg. Technol. Comput. Syst.15
2026 Time-Domain SRAM-CIM Macro With Dual-Edge Temporal Fused Accumulation for Signed 8-bit Precision MAC
Wenjuan Lu, Xiaobo Gong, Kang Meng, Xiaohang Chen, Jiating Guo, Lijun Guan, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu, Chunyu Peng
IEEE Trans. Circuits Syst. I Regul. Pap.10
2026 A T8T-SRAM Computing-in-Memory Macro for Ternary Deep Neural Networks and Boolean Logic Computations
abstract
Deep neural networks (DNNs) play important roles in artificial intelligence applications and show hungry computility and power demands. Compared with binary neural networks (BNNs), ternary neural networks (TNNs) have higher representation and adaptive abilities and balance the inference accuracy and computing efficiency between DNNs and BNNs. This article proposed a T8T-SRAM computing-in-memory (CIM) macro to achieve Boolean logic operations and MAC operation of ternary activation and ternary weight. The proposed T8T-SRAM bitcell has a separate read and write path, and can avoid the read disturb issue. In Boolean logic operation mode, the T8T-SRAM macro can achievenand,nor,xnor, andxoroperations with redundant rows, reducing the additional reference voltage generation circuit. In the MAC mode, the result is quantized by an embedded column analog-to-digital converter (ADC), which uses activation refresh to reduce weight changing. In 28-nm CMOS technology, under 0.5-V array supply voltage and 0.9-V peripheral supply voltage, simulation results manifest that the MAC results have good linearity, and feasibility of Boolean logic operation. The proposed T8T-SRAM macro realizes MAC operation of 16 ternary activations and 16 ternary weights with 333.99–816.1-TOPS/W energy efficiency and 61.9-TOPS/mm2area efficiency. Using an ResNet-18 network for the inference of MNIST, and CIFAR-10 datasets, the accuracies were 99.06% and 85.76% with a ternary activation and ternary weight.
Chenghu Dai, Zihua Ren, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.10
2026 A 28 nm 1.3 TFLOPS/mm2 Floating-Point SRAM-Based CIM Macro With Asynchronous Normalization and Parallel Sorting Alignment for AI-Edge Chip
abstract
State-of-the-art AI edge devices require floating-point (FP) multiply-accumulate (MAC) operations with high-energy efficiency and inference accuracy. FP computing in-memory (FP-CIM) has a broader range of applications compared to integer CIM. However, FP-CIM can incur greater power, delay, and area overheads than integer CIM due to the inherent complexity of FP computational flow. In this article, we introduce a new method for asynchronous exponent normalization and parallel mantissa alignment. This approach allows us to add exponents and find the maximum sum simultaneously. We also replace the traditional subtraction and shifting for mantissa alignment with a cross-structure maximum-finding method, enabling FP-CIM to be achieved with lower delay, area, and power overheads. The macro is designed in TSMC 28 nm process, with a memory size of 6 Kb, a layout area of 0.067 mm2, and an area efficiency of 1.3 TFLOPS/mm2. Simulation results show that the macro computational frequency and energy efficiency can reach 150 MHz and 12.8 TFLOPS/W, respectively, at 900 mV, while performing FP- MAC operations.
Zhi-Ting Lin, Miao Long, Yang Yang 0025, Lintao Chen, Yu Liu 0113, Xin Li 0099, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.10
2026 A Digital FP CIM Macro With Cascaded Row-Elimination for Exponent Comparison and In-Memory Mantissa Sparsity Detection
abstract
Floating-point (FP) computing-in-memory (CIM) addresses the energy efficiency bottleneck of von Neumann architectures and fixed-point CIM in high-accuracy neural network training/inference. However, existing FP CIM designs still suffer from limited exponent-path parallelism and low mantissa-path energy efficiency. This article proposes a novel FP CIM architecture using: 1) a cascaded row-elimination comparison mechanism for single-cycle, high-parallelism exponent max-value comparison; 2) an in-memory sparse feature detection method that skips redundant mantissa computing modules based on different input/weight mantissa patterns to reduce energy consumption; and 3) separated exponent/mantissa CIM modules enabling pipelining and a mantissa bit-extension strategy optimizing accuracy, area, and energy efficiency. Simulation results of a 28-nm 19-kb macro show a 496-MHz operating frequency at 0.9 V and a peak energy efficiency of 12.88 TFLOPS/W.
Wenjuan Lu, Kang Meng, Xiaobo Gong, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.10
2026 A 28-nm 9-kb SRAM Computing-in-Memory Macro With Segmented Charge Sharing for Multimode MAC Operations
abstract
Computing-in-memory (CIM) is a novel approach to solve the von Neumann bottleneck and improve energy efficiency and throughput. This article presents an SRAM-CIM macro based on segmented charge sharing to support multimode multiply-and-accumulate (MAC) operations, including binary weight network (BWN) MAC, ternary weight network (TWN) MAC, and multibit MAC operations. Signed 9-bit input MAC operations can be supported in BWN and TWN networks through multiple cycles. The multibit MAC operations are realized by the cooperation of two cell structures, which eliminates the need to convert negative numbers into complements and reduces the area and power consumption of the chip. Additionally, the proposed segmented charge-sharing scheme increases the speed of accumulation and improves the overall energy efficiency of the chip. The proposed 9-kb macro is implemented in 28-nm CMOS technology with an energy efficiency of 30.74 TOPS/W and an area efficiency of 1.15 TOPS/mm2in multibit MAC mode. At the system level, a ResNet-based model achieves an accuracy of 94.59% on the CIFAR-10 dataset and 74.64% on the CIFAR-100 dataset, demonstrating the effectiveness of the proposed CIM architecture for practical neural network inference. Compared to prior designs, it not only supports additional computing modes but also demonstrates notable improvements in both energy and area efficiency.
Wenjuan Lu, Lening Tan, Tianchen Xue, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.7
2026 A 28-nm 16-Kb SRAM Computing-in-Memory Macro With Dual Bitline Computing for Boolean Logic Operation and BWN MAC Operation
abstract
By integrating computation directly within memory structures, computing-in-memory (CIM) effectively mitigates the von Neumann bottleneck, offering significant improvements in energy efficiency and system throughput. This letter presents a 16-kb static random access memory (SRAM) CIM macro in 28-nm complementary metal–oxide–semiconductor (CMOS), featuring a novel dual bitline computing cell (DBCC) architecture. The design supports three operational modes: conventional memory access, Boolean logic operations, and binary weight network (BWN) multiply-and-accumulate (MAC) computations. The DBCC architecture addresses key critical limitations of prior works by: 1) storing both positive and negative weights within a single 6T-SRAM cell, improving array utilization by$2\times $compared to segregated storage schemes; 2) achieving consistent discharge rates for both polarities through symmetric pull-down paths, enhancing calculation accuracy [INL <3 least significant bit (LSB) across corners] and process variation tolerance; and 3) enabling differential voltage-based Boolean logic operations without external reference voltage generation, thereby reducing peripheral circuitry. The design features a compact 2:1 multiplexer (2:1 MUX)-based pulsewidth modulator (PWM) for 5-b signed input encoding and demonstrates 55.8–204.8-TOPS/W energy efficiency in 28-nm CMOS. Experimental results on a 16-kb array show stable operation across PVT variations while supporting both memory functions and in situ computation for BWNs. System-level evaluation demonstrates its practicality, achieving competitive inference accuracies of 97.59% on MNIST and 89.08% on CIFAR-10.
Wenjuan Lu, Lubin Xiang, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.7
2026 A Floating-Point SRAM Computing-in-Memory Macro Using Digital-Domain Structure for CNNs
Wenjuan Lu, Xiaobo Gong, Xiulong Wu, Chunyu Peng
IEEE Trans. Very Large Scale Integr. Syst.7
2026 A Capacitor Discharge-Based SRAM CIM Macro Based on Hybrid-Domain for Convolutional Neural Networks
abstract
Compute-in-memory (CIM) is increasingly recognized as an effective hardware accelerator for convolutional neural networks (CNNs). This work proposes a hybrid-domain CIM design using: 1) a multibit compute unit (MBCU) structure that realizes the multiplication operation of 2-bit input and 4-bit weight through the transistor-size-weighted capacitor discharge on the bitline; 2) a hybrid-domain quantization scheme (HDQS) of “time-domain + voltage-domain,” which integrates the high energy efficiency of time-domain quantization with the low-delay advantages of the voltage-domain quantization, and enhances the quantization accuracy through the combined effect of the process tracking module and the reference signal module; 3) the CIM circuit design, layout drawing and simulation verification of hybrid-domain static random access memory (SRAM) were realized by 28-nm CMOS technology, results show that the circuit supports 8-bit multiply–accumulate (MAC) operation, and full-precision quantization in the hybrid-domain form can achieve the optimal energy efficiency of 249.7 TOPS/W per bit at 0.7 V, and area efficiency of 4.29 TOPS/mm2per bit. Furthermore, the integration of the circuits with the VGG-16 network has been demonstrated to yield an inference accuracy of 90.52% in the CIFAR-10 dataset.
Bin Qiang, Yongliang Zhou, Xiulong Wu, Chunyu Peng
IEEE Trans. Very Large Scale Integr. Syst.4
2026 A CIM Macro Embedded With Sign Operations for Parallel Signed Multibit Multiplication-and-Accumulation Using Hybrid Cell Array
Jin Zhang 0036, Zhongzhen Tong, Qiang Zhao 0007, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.9
2025 A Floating-Point SRAM-based CIM Macro with Asynchronous Normalization and Parallel Sorting Alignment
abstract
Floating-point computing-in-memory (FP-CIM) has a broader range of applications compared to integer CIM. However, FP-CIM can incur greater power, delay, and area overheads than integer CIM due to a more complex computational flow. In this paper, we introduce a new method for asynchronous exponent normalization and parallel mantissa alignment. This approach allows us to add exponents and find the maximum sum simultaneously. We also replace the traditional subtraction and shifting for mantissa alignment with a time-cycle lookup method, enabling FP-CIM to be achieved with lower delay, area, and power overheads. The macro is designed in the TSMC 28nm process, with a memory size of 6Kb, a layout area of 0.067mm2, and an area efficiency of 1.3TFLOPS/mm2. Simulation results show that the macro computational frequency and energy efficiency can reach 150MHz and 12.8TFLOPS/W, respectively at 900mV.
Zhi-Ting Lin, Dongcheng Wang, Rongtao Li, Shichen Yu, Yu Liu 0113, Xin Li 0099, Xiulong Wu
ISCAS10
2025 TSCIM: A 28nm Transposed Stochastic CIM Macro for On-Chip Training and Inference
abstract
This work introduces a novel Transposed Stochastic Computing-in-Memory (TSCIM) macro designed to enhance the efficiency of on-chip training and inference. The macro incorporates a novel stochastic quantization strategy and utilizes a transposed separated wordline SRAM to enable multi-bit signed MAC operations. Furthermore, a stochastic adder tree is utilized to minimize area and power consumption overhead. The design includes a 4Kb SRAM CIM macro implemented in 28 nm CMOS technology. Simulation results show that the power consumption of the stochastic accumulation circuit (SAC) is reduced by 63.6%, while the area overhead is decreased by a factor of 7.73 compared to designs using full adder (FA) adder trees. Additionally, the computation latency is decreased by 16× compared to traditional stochastic circuits. The TSCIM macro can achieve a peak energy efficiency of 63.02 TOPS/W and an area efficiency of 15.54 TOPS/mm2.
Yu Liu 0113, Yang Lou, Kangkang Mao, Xin Li 0099, Chenghu Dai, Xiulong Wu, Zhi-Ting Lin
ISCAS6
2025 MTJ based temperature compensated beta multiplier Voltage Reference
abstract
This article mainly explores the interaction between CMOS devices and Magnetic Tunnel Junction (MTJ) devices in terms of temperature characteristics, aiming to achieve a CMOS beta multiplier circuit that combines low power consumption and wide temperature adaptability, making it a stable reference voltage source. The proposed design utilizes the Tunneling Magneto Resistance (TMR) effect of MTJ to compensate for the performance mismatch caused by temperature changes in CMOS. This design adopts TSMC 28nm CMOS craft, which can generate a reference voltage with a linearity of 0.57%/V, a temperature coefficient of 43.6ppm/°C, and a stable voltage of 441.6mV at a minimum supply voltage of 0.6V and a temperature range of 10~110 °C. The noise of this voltage at a frequency of 10Hz is 69.7uV/sqrt (Hz), the power suppression ratio is -62.1dB, and the power consumption is 4.092nW.
Yongliang Zhou, Yingxue Sun, Jingxue Zhong, Chengxing Dai, Weizhe Tan, Chunyu Peng, Xin Li 0099, Zhi-Ting Lin, Xiulong Wu
ISCAS10
2025 MTJ based Temperature-Adaptive VCO (TAVCO) for Compensating CP-PLL Frequency Drift
abstract
The Charge Pump Phase-Locked Loop (CP-PLL) is a commonly utilized component in contemporary mixed-signal electronic systems. It is widely employed for clock generation, synchronization, and frequency synthesis in both digital and wireless functionalities. However, the frequency accuracy of oscillators can be adversely affected by variations in frequency across a broad temperature range. To address this issue, the voltage-controlled oscillator designed in this study employs a four-stage differential delay structure, chosen for its simple circuit architecture, favorable control linearity, and low noise characteristics. This research integrates the temperature behaviors of Complementary Metal-Oxide-Semiconductor (CMOS) and Magnetic Tunnel Junction (MTJ) technologies, utilizing 28nm CMOS technology to enhance the frequency stability of ring oscillators effectively. Simulation results indicate that frequency drift is reduced by 92% within the temperature range of -80°C to 125°C.
Yongliang Zhou, Jingxue Zhong, Yingxue Sun, Chengxing Dai, Weizhe Tan, Chunyu Peng, Wenjuan Lu, Xin Li 0099, Zhi-Ting Lin, Xiulong Wu
ISCAS10
2025 An RISC-V PPA-Fusion Cooperative Optimization Framework Based on Hybrid Strategies
abstract
The optimization of RISC-V designs, encompassing both microarchitecture and CAD tool parameters, is a great challenge due to an extensive and high-dimensional search space. Conventional optimization methods, such as case-specific approaches and black-box optimization approaches, often fall short of addressing the diverse and complex nature of RISC-V designs. To achieve optimal results across various RISC-V designs, we propose the cooperative optimization framework (COF) that integrates multiple black-box optimizers, each specializing in different optimization problems. The COF introduces the landscape knowledge exchange mechanism (LKEM) to direct the optimizers to share their knowledge of the optimization problem. Moreover, the COF employs the dynamic computational resource allocation (DCRA) strategies to dynamically allocate computational resources to the optimizers. The DCRA strategies are guided by the optimizer efficiency evaluation (OEE) mechanism and a time series forecasting (TSF) model. The OEE provides real-time performance evaluations. The TSF model forecasts the optimization progress made by the optimizers, given the allocated computational resources. In our experiments, the COF reduced the cycle per instruction (CPI) of the Berkeley out-of-order machine (BOOM) by 15.36% and the power of Rocket-Chip by 12.84% without constraint violation compared to the respective initial designs.
Tianning Gao, Ming Zhu 0016, Xiulong Wu, Dian Zhou, Zhaori Bi
IEEE Trans. Very Large Scale Integr. Syst.4
2025 A Low-Cost and Triple-Node-Upset Self-Recoverable Latch Design With Low Soft Error Rate
abstract
With the decrease in feature size of transistors, latches are more sensitive to single-event multiple node upset (MNU), including double node upset (DNU) and triple node upset (TNU). However, the reported TNU self-recoverable (TNUR) latches are facing problems with large areas and power consumption. Based on the polarity design, this article proposes a low-cost TNUR latch (LCTRL) with a low soft error rate (SER) in 28-nm CMOS technology. The proposed LCTRL mainly consists of four interlocked modules and a clock-gated inverter. Compared with the state-of-the-art TNUR latches, including LCTNURL, IHTRL, FATNU, and TRLW, the power consumption, D-Q delay, CLK-to-Q delay, area, and the power-delay–area product (PDAP) of the proposed LCTRL are reduced by 55.09%, 38.64%, 42.93%, 44.65%, and 83.50%, respectively. Due to the polarity design, the SER of the proposed LCTRL is the smallest among compared latches, which suggests that the proposed LCTRL is suitable for use in radiation environments.
Licai Hao, Lang Tian, Hao Wang 0239, Shiyu Zhao 0004, Qiang Zhao 0007, Chunyu Peng, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.9
2025 Full-Array Boolean Logic CIM Macro With Self-Recycling 10T-SRAM Cell for AES Systems
abstract
Computing in memory (CIM), which alleviates the need to transfer a large amount of data between processor and memory, significantly reducing latency and energy consumption, is a promising new computing architecture for addressing the von Neumann bottleneck problem. This article proposes a CIM array structure composed of self-recycling 10T static random access memory (SRAM) cells, which can realize orthogonal data writing, and multiple Boolean logical operations for the entire array. The self-recycling and full-array activation characteristics are extremely suitable for accelerating diverse data processing algorithms such as the Advanced Encryption Standard (AES). A 4-kb SRAM is implemented in 55-nm CMOS technology to verify the effectiveness of the design. Compared with other state-of-the-art architectures, the throughput and the operating frequency of the proposed CIM macro are increased to 843 GOPS/kb ($2.64\times $) and 823.7 MHz ($2.6\times $), respectively. The energy efficiency reaches 246.9 TOPS/W. When applied to the AES, the energy consumption is 35.77% less than the digital CIM architecture that is not self-recycling.
Xin Li 0099, Lintao Chen, Yang Lou, Baofa Wu, Jiajun Long, Yongliang Zhou, Chunyu Peng, Xiulong Wu, Zhi-Ting Lin
IEEE Trans. Very Large Scale Integr. Syst.10
2025 A 28-nm 9T1C SRAM-Based CIM Macro With Hierarchical Capacitance Weighting and Two-Step Capacitive Comparison ADCs for CNNs
abstract
In the realm of charge-domain computing-in-memory (CIM) macros, reducing the area of capacitor ladder and analog-to-digital converter (ADC) while maintaining high throughput remains a significant challenge. This brief introduces an adjustable-weight CIM macro designed to enhance both energy efficiency and area efficiency for convolutional neural networks (CNNs). The proposed architecture uses: 1) a customized 9T1C bit cell for sensing margin improvement and bidirectional decoupled read ports; 2) a hierarchical capacitance weighting (HCW) structure that achieves a weight accumulation of 1/2/4 bits with less capacitance area and weighting time; and 3) a two-step capacitive comparison ADCs (TC-ADCs) readout scheme to improve area efficiency and throughput. The proposed 8-kb static random address memory (SRAM) CIM macro is implemented using 28-nm CMOS technology. It can achieve an energy efficiency of 224.4 TOPS/W and an area efficiency of 21.894 TOPS/mm2, and the accuracies on MNIST, CIFAR-10, and CIFAR-100 datasets are 99.67%, 89.13%, and 67.58% with a 4-b input and 4-b weight.
Zhi-Ting Lin, Runru Yu, Miao Long, Yu Liu 0113, Jianxing Zhou, Qingchuan Zhu, Yue Zhao 0029, Lintao Chen, Chunyu Peng, Qiang Zhao 0007, Xin Li 0099, Chenghu Dai, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.15
2025 High-Reliability and High-Throughput CIM 10T-SRAM for Multiplication and Accumulation Operations With 274.3 GOPS and 200-237.5 TOPS/W
abstract
Artificial intelligence (AI) is extensively applied in natural language processing, image matching, and image recognition, with convolutional neural networks (CNNs) being crucial. Computing-in-memory (CIM) utilizing static random access memory (SRAM) can enhance the CNN performance. However, this faces issues such as multibit signed data processing, read corruption of traditional SRAM arrays, and increased area overhead due to increased capacitor weighting. This article proposes a 10T-SRAM macro tailored for CNN multiply-accumulate calculation (MAC) computation in image processing. It enables high-throughput full-array operations, with added dual ports facilitating input of multibit data with signed bits. The 10T-SRAM cell features a read-write separation channel, mitigating read disturbance issues seen in dual-port 8T-SRAM arrays or 6T-SRAM arrays. Incorporating redundant columns in the array for charge sharing and weighting conserves area and boosts circuit reliability. In the 28-nm CMOS simulation environment, the proposed architecture achieves a throughput of 274.3 GOPS and an energy efficiency of 200–237.5 TOPS/W, surpassing literature-reported figures by several times.
Wenjuan Lu, Lubin Xiang, Chunyu Peng, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.7
2025 A 28 nm Dual-Mode SRAM-CIM Macro With Local Computing Cell for CNNs and Grayscale Edge Detection
abstract
With the rise of artificial intelligence (AI), neural network applications are growing in demand for efficient data transmission. The traditional von Neumann architecture can no longer keep pace with modern technological needs. Computing-in-memory (CIM) is proposed as a promising solution to address this bottleneck. This work introduces a local computing cell (LCC) scheme based on compact 6T-SRAM cells. The proposed circuit aims to enhance energy efficiency and reduce power consumption by reusing the LCC. The LCC circuit can perform the multiplication of a 2-bit input with a 1-bit weight, which can be applied to convolutional neural networks (CNNs) with the multiply-accumulate (MAC) operations. Through circuit reuse, it can also be used for multibit multiply operations, performing 2-bit input multiplication and 1-bit weight addition, which can be applied to grayscale edge detection in images. The energy efficiency of the SRAM-CIM macro achieves an energy efficiency of 46.3 TOPS/W under MAC operations with input precision of 8-bits and weight precision of 8-bits, and up to 389.1–529.1 TOPS/W under the calculation in one subarray with an input precision of 2-bits and a weight precision of 1-bit. The estimated inference accuracy on CIFAR-10 datasets is 90.21%.
Chunyu Peng, Xiaohang Chen, Mengya Gao, Jiating Guo, Lijun Guan, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.8
2025 A 28-nm Cascode Current Mirror-Based Inconsistency-Free Charging-and-Discharging SRAM-CIM Macro for High-Efficient Convolutional Neural Networks
abstract
Computing-in-memory (CIM) is an emerging approach to alleviate the von Neumann bottleneck and enhance energy efficiency and throughput. This brief introduces a 16-Kb static random access memory (SRAM) CIM macro for convolutional neural networks (CNNs), featuring a cascode current mirror-based inconsistency-free computing circuits (CICCs). The bias voltage of CICC is provided by a cascode current mirror (CCM) circuit. The proposed architecture improves the consistency and linearity of bitline (BL) charge and discharge rates in the analog current domain, enhancing computational accuracy. Additionally, the charge and discharge on the BLs represent the positive or negative calculation result, eliminating the need for extra encoding and logic circuits to handle sign bits. The SRAM-CIM macro achieves an energy efficiency of 59.1–134.0 TOPS/W and a throughput of 0.41 TOPS in a 28-nm CMOS technology, and the estimated inference accuracy on MNIST and CIFAR-10 datasets is 96.5% and 91.4%, respectively, with 5-bit input precision and 1-bit weight precision.
Chunyu Peng, Jiating Guo, Shengyuan Yan, Xiaohang Chen, Wenjuan Lu, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.9
2025 A High-Performance and High-Robustness Triple-Node-Upset Tolerant Latch Based on Redundant-Node Hardening
abstract
In response to the issues of high cost, large overhead, and limited node fault tolerance in current latch hardening techniques, this article proposes a latch circuit resistant to triple-node-upset (TNU) based on redundant-node hardening technology. This latch comprises eight 1P2N modules interlocked, with its output isolated by two levels of C-elements (CEs), achieving tolerance to TNU. The performance of the redundant-node reinforcement TNU tolerant latch (RNRTTL) was simulated and verified using CMOS 65 nm technology. The simulation results indicate that the RNRTTL circuit has a D-Q delay of 14.14 ps, static power consumption of$4.03~\mu $w, an area of$32.87~\mu $m2, and an area-static power-D–Q delay-product (APDP) of 1873, respectively. Compared to the triple-node upset tolerant latches TTLL, TNU-latch, TNURL, and HLTNURL reported in the current literature, the proposed latch demonstrates an average reduction of 219.9%, 164.9%, 150.7%, and 2464.8% in D-Q delay, static power consumption, area, and APDP, respectively, indicating that the RNRTTL latch has superior comprehensive performance; furthermore, a series of 2000 Monte Carlo (MC) simulations on the node group$\langle $Q, X0, X$8\rangle $reveal that the proposed latch circuit possesses good stability, making it suitable for harsh radiation environments.
Qiang Zhao 0007, Qingyi Liu, Licai Hao, Xin Li 0099, Shengyue Zhang, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.9
2024 SRAM-Based Digital CIM Macro for Linear Interpolation and MAC
abstract
Linear interpolation is widely used in algorithms such as image segmentation, but the existing compute-in-memory (CIM) architectures cannot satisfy the needs of linear interpolation. This paper proposes a CIM macro based on static random-access memory (SRAM) that implements linear interpolation for the first time. Swing multiplication and accumulation are proposed in this paper for linear interpolation operations. In addition, the proposed circuit can be used for multiply-and-accumulate (MAC) operation and supports parallel updating and computing. A new adder tree by reused the adders inside the multiplier to implement the accumulation operation. The proposed circuit can improve the area efficiency as no additional adder tree circuit is required. The design is implemented in a 28 nm process and area efficiency can achieve 0.059 TOPS/mm2for MAC operation. When operating with an 8-bit linear interpolation, this CIM macro achieves access times of 6−9 ns and energy efficiencies of 11.13−17.72 TOPS/W. Under MAC operation with 8-bit inputs and weights, this CIM macro achieves access time of 6.36−9.47 ns and energy efficiency of 4.61−8.36 TOPS/W for 19-bit outputs.
Zhi-Ting Lin, Yunlong Liu 0006, Yaling Wang, Yue Zhao 0029, Chunyu Peng, Xiulong Wu
ISCAS6
2024 A Timing-Shared Adaptive Sensing Methodology for Low-Voltage SRAM
abstract
Lowing static random access memory (SRAM) supply voltage could highly improve energy efficiency, yet energy efficiency still not attains optimal point due to the constraint of the weakest bit-cell, especially in low-voltage SRAM. Adaptive sensing methodology is proposed for the challenge and consists of four elements: Switch unit is built to implement cross-sensing operation, Timing-Shared Decoupling Latch Sense Amplifier (TS-DLSA) allows rapid and successive sensing, the judging module is utilized to trigger the FLAG signal which is for adaptive timing controller to cut off word line (WL). Energy efficiency was obtained by compressing the activation delay of WL compared to global timing scheme. The proposed adaptive sensing methodology was performed with TSMC 28-nm CMOS process, and evaluated in SRAM array of 128x128, 256x256, 512x512, and 1024x1024. Monte Carlo simulation results are formed to confirm that, compared to the global timing scheme, the proposed adaptive sensing methodology has reduced WL activation delay by 75.1%∼41.8% and read operation energy overhead by 77.5%∼30.9% from 0.6V to 1.2V. Compared to the current-latched sense amplifier with a footswitch (FS-CLSA) with proposed adaptive sensing methodology, TS-DLSA with proposed scheme have reduced the WL activation by 4.1%∼10.5% and read operation energy overhead by 15.4%∼23.8% from 256x256 to 2048x2048 at 0.6V. The more cells mounted on the BL, the higher energy revenue gains.
Yongliang Zhou, Saiai Wu, Wenjuan Lu, Chunyu Peng, Xin Li 0099, Xiulong Wu
ISCAS9
2024 A High Throughput In-MRAM-Computing Scheme Using Hybrid p-SOT-MTJ/GAA-CNTFET
abstract
Silicon-based semiconductor transistors are approaching their physical limits due to shrinking feature sizes. Simultaneously, traditional silicon-based von Neumann architectures exhibit significant latency and power consumption issues in data-centric applications, such as the Internet of Things and artificial intelligence. To tackle these challenges, this study introduces a novel approach: Magnetoresistance Random Access Memory (MRAM) computing in-memory (CIM) using gate-all-around carbon nanotube field-effect transistors (GAA-CNTFET). The proposed MRAM array comprised three transistors and one perpendicular magnetic anisotropy spin-orbit torque magnetic tunnel junction (p-SOT-MTJ) (3T1M) cell and achieves full-array Boolean logic operations and half/full-adder operations. The calculated results can be stored in-situ during the computing phase without requiring additional peripheral circuits. A 16 Kb MRAM was simulated in both GAA-CNTFET/p-SOT-MTJ and 14-nm FinFET/p-SOT-MTJ technologies to examine the effectiveness of the proposed design. Compared to its 14-nm FinFET/p-SOT-MTJ counterparts, the write and computing latencies of the GAA-CNTFET/p-SOT-MTJ CIM macro were reduced by approximately 21% and 20.6%, respectively, while the read and computing energy consumption by approximately 45.3% and 24.7%, respectively. Moreover, the proposed in-memory Boolean logic throughput was 8192 GOPS, which was approximately 160–250 times higher than that of existing CIM solutions, in which only two rows of word lines can be activated.
Zhongzhen Tong, Yunlong Liu 0006, Xinrui Duan, Suteng Zhao, Chenghang Li, Zhi-Ting Lin, Xiulong Wu, Zhaohao Wang, Xiaoyang Lin
IEEE Trans. Circuits Syst. I Regul. Pap.9
2024 A Computing In-Memory Multibit Multiplication Based on Decoupling and In-Array Storing
abstract
Multiplications are basic operations of neural networks. Therefore, multiplication results are crucial in analyzing the operating process of neural networks. However, the multiplication strategies are generally based on analog-domain circuits, and the results are in a multiply-and-accumulate (MAC) form. The result of each multiplication in MAC cannot be distinguished accurately using these strategies. Therefore, we proposed an in-memory multibit multiplication based on the decoupling and in-array storage strategy to overcome this problem, and the core module is the 10T1C SRAM cell. Multibit multiplications are decoupled by a series of logical operations. Therefore, in the analysis mode, multiplication results can be saved and outputted in the normal read mode without requiring additional storage. When executing the neural network, the operation results are stored in the cells. Hence, the operands stored in the array are retained. Accumulation operations are completed based on the charge-sharing technology; thus, the linearity of accumulation is high. We simulated and analyzed the performance of the proposed circuit in a 28 nm CMOS process. The absolute value of integral nonlinearity is at most 0.29. Further, due to high data operation parallelism, the throughputs of the logical operation and MAC are up to 6307.8 and 802.8 GOPS, respectively.
Jin Zhang 0036, Zhongzhen Tong, Hao Wang 0239, Qiang Zhao 0007, Jiaqun Wang, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Circuits Syst. I Regul. Pap.9
2024 Timing Optimization Model and PVT Tracked Scheme for STT-MRAM Voltage-Mode Sense
abstract
The impact of process variations on the read operation of low-voltage STT-MRAM becomes severe, posing a challenge in determining the optimal sensing timing of the sense amplifier. This study investigates techniques for refining the timing scheme of sensing circuits in order to improve the sensing reliability of the STT-MRAM. The supply voltage$V_{DD}$, the Tunneling Magnetoresistance Ratio TMR, the low resistance state of bit-cell$R_{P}$, and the parasitic capacitance of bit-line$C_{BL}$are analyzed along with the voltage sense amplifier (VSA) involved in sensing yield. We develop a timing model through theoretical analysis to determine the optimal VSA enable signal (SAE). In addition, an innovative Process-Voltage-Temperature (PVT) tracking scheme is proposed that can track the optimal VSA enable signal (SAE) and suppress timing variations. Monte-Carol simulation in the 28nm CMOS and magnetic tunnel junction (MTJ) process confirms that the combined scheme significantly enhances the robustness of sensing operation. The proposed scheme improves yield by 20% to 35%, reduces power consumption by 43% to 63%, and reduces read access delay by 47% to 59% compared to conventional sensing schemes at 0.6V supply voltage.
Yongliang Zhou, Yingxue Sun, Chengxing Dai, Jingxue Zhong, Xiulong Wu, Chunyu Peng
IEEE Trans. Circuits Syst. I Regul. Pap.9
2024 A CFMB STT-MRAM-Based Computing-in-Memory Proposal With Cascade Computing Unit for Edge AI Devices
abstract
The application of non-volatile memory technology is increasingly attractive for Computing-in-memory (CIM) owing to high integration density and negligible standby power consumption. This study proposes an spin-transfer-torque (STT) magnetic random access memory (MRAM) based CIM macro which incorporates following innovative features: 1) cross-feedback margin-boost (CFMB) scheme to enable robust and fast reading operations against process variation and limited Tunneling Magnetoresistance Ratio (TMR); 2) cascade computing units (CCU) and related design method for efficient and stable multi-bit multiply-and-accumulate (MAC) operation; and 3) dual computing mode scheme and resolution adjustable quantization module to optimize energy efficiency and operating speed. The post-simulations are performed under 28nm CMOS&MTJ technology. The results demonstrate the achievement in energy efficiency of 36.4 TOPS/W while performing MAC operations with up to 16-bit weights, 4-bit inputs, and 22-bit outputs.
Yongliang Zhou, Chenghu Dai, Licai Hao, Chunyu Peng, Hao Cai 0001, Xiulong Wu
IEEE Trans. Circuits Syst. I Regul. Pap.10
2024 Low-Cost and Highly Robust Quadruple Node Upset Tolerant Latch Design
abstract
This article proposes an exceptionally reliable and low-cost quadruple node upset tolerant latch ($LC$-QNUTL) suitable for the 65 nm CMOS technology. The innovative$LC$-QNUTL latch is primarily composed of three soft-error-immune (SEI) static random-access memory (SRAM) cells and a triple-level C-element (CE) unit, which includes five two-input CE and a clock-gating (CG)-based two-input CE. The SEI SRAM cell utilizes polarity hardening technology and source-isolation technology, significantly reducing the number of sensitive nodes and enhancing the latch’s stability. By using the high-speed transmission gate (TG) technology and stacked structures, the proposed latch offers minimal overhead in terms of delay and power consumption, yielding an improved power delay area product (PDAP). When compared to contemporary quadruple node upset (QNU)-tolerant latch designs (including HLMR, 4NUHL, and LDAVPM), the new design offers substantial improvements—29.53% less delay, 80.09% reduced power consumption, 58.52% smaller silicon area, and 433.43% improved comprehensive PDAP on average. Furthermore, simulation results demonstrate that the$LC$-QNUTL latch exhibits reduced sensitivity to process, voltage, and temperature (PVT) variations, thus providing superior reliability, which makes it an ideal choice for safety-critical applications.
Licai Hao, Yaling Wang, Yunlong Liu 0006, Shiyu Zhao 0004, Wenjuan Lu, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.13
2024 Soft-Error-Immune Quadruple-Node-Upset Tolerant Latch Based on Polarity Design and Source-Isolation Technologies
abstract
A soft-error-immune quadruple-node-upset tolerant latch (SEI-QNUTL) with a low delay and high performance is proposed using 65-nm CMOS technology. The proposed SEI-QNUTL design consists of three soft-error-immune static random access memory (SEI-SRAM) cells. Furthermore, each SEI-SRAM cell employs polarity design and source-isolation technology to reduce the number of sensitive nodes and enhance the reliability of the latch. Compared with state-of-the-art quadruple-node-upset (QNU) tolerant latches [including high-performance and low-cost single-event multiple-node-upsets resilient (HLMR), QNU tolerant latch (QNUTL), and Latch Design and Algorithm-based Verification Protected against Multiple-Node-Upsets (LDAVPM)], the proposed SEI-QNUTL design reduces (on average) the area, delay, and area-power-delay-product (APDP) by 47.0%, 25.0%, 46.5%, and 66.3%, respectively. Extensive variation analysis validates that the SEI-QNUTL design is less sensitive to process, voltage, and temperature (PVT) variations regarding power consumption and delay. Furthermore, Monte Carlo (MC) simulations show that the proposed latch exhibits high reliability when performing data storage. Compared with the existing latches, the SEI-QNUTL design makes a good tradeoff among delay, power, and area, and it can thus be used in safety-critical applications.
Licai Hao, Chenghu Dai, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng, Yongliang Zhou, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.9
2023 High Restore Yield NVSRAM Structures With Dual Complementary RRAM Devices for High-Speed Applications
abstract
Static random access memory (SRAM) plays a key role in the overall performance of electronic systems because of its rapid data processing and transmission speed; however, when the system power supply is cut off, the data stored in the nodes are lost. Thus, this article proposes four nonvolatile SRAM (NVSRAM) cells that use unilateral or bilateral structures with dual complementary series resistive random access memory (RRAM) devices. It is found that the read, write, and hold static noise margins (HSNMs) are comparable with those of the standard 6T-SRAM. Moreover, the store and restore operations operate in parallel at high speed. The store operation delay is only 6 ns for unilateral structures and 5 ns for bilateral structures, and the restore delay is only 10 ns for unilateral structures and 6 ns for bilateral structures. The maximum power consumption among the four structures for storing and restoring a “1” are 1.545 pJ/bit and 134.5 fJ/bit, respectively. Furthermore, the dual complementary series resistor structures can achieve a high restore yield at a resistance ratio of 1.5. Therefore, a high restore yield can be achieved even with large resistance fluctuations caused by the voltage, time, and process.
Zhi-Ting Lin, Xiulong Wu, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng
IEEE Trans. Very Large Scale Integr. Syst.4
2023 A Fully Digital SRAM-Based Four-Layer In-Memory Computing Unit Achieving Multiplication Operations and Results Store
abstract
The separation of memory and arithmetic logic unit (ALU) in the von Neumann computing architecture hinders the development of big data and high-performance computing. In-memory computing (IMC) as a new computation method significantly reduces the latency and power consumption of data processing. In this study, we propose a fully digital static random access memory (SRAM)-based IMC architecture, which has the following advantages: 1) it simplifies multiplication to multicycle addition operations, reuses logic cells, and reduces hardware overhead; 2) by adding a pair of nMOS transistors to achieve internal write-back, the computational efficiency is improved, and at the same time, the final result of the multiplication can be stored locally, eliminating the need to read the computational result immediately; and 3) this scheme can be easily expanded to multiplication operations with different bit widths, which provides good scalability. A 4-kb SRAM-IMC macro chip is manufactured using the SMIC 55-nm technology to realize 4-bit multiplication, with an energy efficiency of 51.4 TOPS/W (0.9 V) and a throughput of 234.3 GOPS/mm2. The proposed multiplication–accumulation architecture is applied to a neural network, which achieves 98.7% accuracy with the Mixed National Institute of Standards and Technology database (MNIST) dataset.
Zhi-Ting Lin, Shaoying Zhang, Jianping Xia, Yunwei Liu, Kefeng Yu, Zhongzhen Tong, Xiulong Wu, Wenjuan Lu, Chunyu Peng, Qiang Zhao 0007
IEEE Trans. Very Large Scale Integr. Syst.12
2023 In-Memory Transposable Multibit Multiplication Based on Diagonal Symmetry Weight Block
abstract
A possible approach to overcome the von Neumann bottleneck and meet the increasing demand for better computing performance is to computing in-memory (CIM). The results of the in-memory calculations are primarily reflected in the vertical bitline (BL) analog voltage. However, the nonlinearity of the BL discharge deteriorates with the increase in discharge voltage. In this study, we propose a diagonal symmetry weight block (DSWB) based on an eight-transistor (8T) static random access memory (SRAM) that can achieve multibit transposable operations. In addition, to guarantee linearity and complete multibit multiplication operations, we propose a cascode current mirror (CCM)-based multiplier. To achieve low-overhead and more efficient quantification, our proposed CIM macro uses a counter-type quantization circuit to read out the analog calculation results. We simulated the performance of the proposed 8T SRAM in a 28-nm complementary metal–oxide–semiconductor process. The integral nonlinearity (INL) of the proposed CCM-based CIM decreased by approximately 54.4% compared with the traditional CIM. Furthermore, the proposed in-memory multibit multiplication throughput density was 6.74 GOPS/kb; this throughput density improvement is approximately 3.3–10.5 times higher than the existing CIM works.
Zhongzhen Tong, Yue Zhao 0029, Jin Zhang 0036, Zhi-Ting Lin, Xiaoyang Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.6
2022 Configurable Memory With a Multilevel Shared Structure Enabling In-Memory Computing
abstract
Frequent to-and-from data transfers in the von Neumann architecture limit the overall throughput. One of the promising approaches used to overcome von Neumann bottleneck is in-memory computing (IMC) that aims to embed computing in memory to reduce the transfer of memory-processor data. This study proposes a configurable 6-transistor (6T) static random access memory (SRAM) array with a multilevel shared structure for IMC. A multilevel shared structure can effectively improve the utilization rate of the module. In addition to the conventional SRAM operation, the configurable structure can also perform the sum of absolute differences (SAD) and Hamming distance (HD) calculations. To quickly identify the minimum value among multiple calculation results, a four-input sense amplifier (SA) is proposed. The performance of the proposed memory is simulated in a 65-nm CMOS process. The post-layout simulation results show good linearity of the multirow read in the SAD and HD modes. The mean time required by the four-input SA to obtain the result is 190 ps. The SAD and HD calculations yield consumptions of 67.44 fJ/byte and 0.64 fJ/bit, respectively, at 0.8 V. Furthermore, a single column-sharing comparator consumes 2.78 and 3.41 pJ at 0.8 V in the SAD and HD modes, respectively.
Yue Zhao 0029, Zhi-Ting Lin, Xiulong Wu, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng, Zhongzhen Tong, Junning Chen
IEEE Trans. Very Large Scale Integr. Syst.3
2020 An Efficient and Robust Yield Optimization Method for High-dimensional SRAM Circuits
abstract
Due to time-consuming SPICE simulations and extremely low failure rates, yield optimization for large static random access memory (SRAM) circuits is still a challenging problem. In this paper, a novel robust yield optimization problem is firstly proposed for SRAM circuits, where robust means considering design and process parameter variations simultaneously. Both a multi-fidelity Gaussian process regression model, which utilizes the strong nonlinear relationship between small and large SRAM columns, and a Bayesian optimization framework are applied to guide the sampling of the expensive large SRAM circuits. A multimodal problem is formulated to find all peaks and valleys on the small SRAM circuits. Such precomputational knowledge can accelerate the convergence of the proposed multi-fidelity and Bayesian optimization framework. Experimental results show that robust yield is essential to yield optimization, for traditional optimal design will degenerate with 4-5 orders of magnitude of yields, if design variations considered, and it doesn't coincide with the new optimum under the robust yield. The proposed method can gain a 3~4× speedup compared to the state-of-the-art method without loss of accuracy.
Tianchen Gu, Changhao Yan, Xiulong Wu, Fan Yang 0001, Sheng-Guo Wang, Dian Zhou, Xuan Zeng 0001
DAC4
2020 Multiple Sharing 7T1R Nonvolatile SRAM With an Improved Read/Write Margin and Reliable Restore Yield
abstract
This article presents a new resistive random access memory (RRAM)-based average 7T1R nonvolatile static random access memory (nvSRAM). This multiple sharing (MS) 7T1R uses MS schemes in which some of the transistors play various roles. Therefore, the MS-7T1R can perform a decoupled read, enhance write capability, and improve the restore yield with a small area cost. Furthermore, the MS-7T1R can offer two alternative modes, i.e., a high speed and a stable mode. Compared with existing technologies, such as the previous 6T-based nvSRAMs, the results show that the proposed architecture provides a remarkable restore yield and ~154% improvement in the read static noise margin (at TT corner and stable mode). In addition, the read delay improves by ~23% (at TT corner and high-speed mode). The write “1” problem of the single bitline is effectively resolved with our proposed write strategy. The static write margin of “1” is improved by ~88.6% compared with the conventional 6T (β = 4) at a power of 1.2 V. In addition, dynamic power is effectively reduced by the use of the single bitline and sub-word-line driver technology.
Zhi-Ting Lin, Yong Wang 0038, Chunyu Peng, Xiulong Wu, Junning Chen
IEEE Trans. Very Large Scale Integr. Syst.4
2020 In-Memory Computing With Double Word Lines and Three Read Ports for Four Operands
abstract
The von Neumann architecture is approaching its limits in terms of scalability and power consumption. In-memory computation is a possible approach to mitigate this limitation. This brief proposes a configurable 8T static random access memory (SRAM) cell with double word lines and three read ports for in-memory computing. In addition to the normal SRAM function, XOR/XNOR and compound Boolean logic operations of three or four operands, such as AND-OR, AND-OR-INVERT, OR-AND, and OR-AND-INVERT, can be performed in one cycle by fully utilizing the three read ports to obtain 13.2-fJ/bit consumption at 0.6 V. The logic operation frequency is 793 MHz at 1.2 V. The proposed SRAM effectively resolves the bottleneck of the existing in-memory computation schemes that only support compound Boolean logic operations with more than two cycles. In addition, the proposed SRAM array scheme can be configured and used as a binary content-addressable memory or a ternary content-addressable memory for searching operations; it achieves 0.24 fJ/search/bit at 0.6 V in the worst case. At 1.2 V, the searching frequency is up to 813 MHz when searching 128 bits with 65-nm technology.
Zhi-Ting Lin, Honglan Zhan, Chunyu Peng, Wenjuan Lu, Xiulong Wu, Junning Chen
IEEE Trans. Very Large Scale Integr. Syst.6
2020 Novel Write-Enhanced and Highly Reliable RHPD-12T SRAM Cells for Space Applications
abstract
In this brief, we proposed, based on the polarity upset mechanism of single-event transient voltage of n-channel metal-oxide-semiconductor (nMOS) transistors, a novel radiation hardened by polar design (RHPD) 12T SRAM cell to enhance the reliability and operation speed for space applications. Simulation results in Semiconductor Manufacturing International Corporation (SMIC) 65-nm CMOS commercial standard process show that the proposed RHPD-12T cell can tolerate all single-node upsets. Meanwhile, compared with We-QUATRO, QUATRO, and dual interlocked storage cell (DICE), the write speed of the proposed cell can be reduced by ~41.8 and ~35.3%, and the static power consumption is reduced by ~41.6 and ~46.3%, respectively. Monte Carlo (MC) simulation has proved that under high frequency and low supply (0.6 V) voltage, RHPD-12T has the minimum write failure probability compared with five other SRAM cells.
Qiang Zhao 0007, Chunyu Peng, Junning Chen, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.5
2019 Radiation-Hardened 14T SRAM Bitcell With Speed and Power Optimized for Space Application
abstract
In this paper, a novel radiation-hardened 14-transistor SRAM bitcell with speed and power optimized [radiation-hardened with speed and power optimized (RSP)-14T] for space application is proposed. By circuit- and layout-level optimization design in a 65-nm CMOS technology, the 3-D TCAD mixed-mode simulation results show that the novel structure is provided with increased resilience to single-event upset as well as single-event-multiple-node upsets due to the charge sharing among OFF-transistors. Moreover, the HSPICE simulation results show that the write speed and power consumption of the proposed RSP-14T are improved by ~65% and ~50%, respectively, compared with those of the radiation hardened design (RHD)-12T memory cell.
Chunyu Peng, Jiati Huang, Changyong Liu, Qiang Zhao 0007, Songsong Xiao, Xiulong Wu, Zhi-Ting Lin, Junning Chen, Xuan Zeng 0001
IEEE Trans. Very Large Scale Integr. Syst.6
2018 Average 7T1R Nonvolatile SRAM With R/W Margin Enhanced for Low-Power Application
abstract
A new average 7T1R nonvolatile SRAM for low-power application is presented in this brief, which improves the read and write margin (RM/WM), as well as the restore energy, simply by using the source switch transistor. Simulation results demonstrate that the RM and WM will be improved by ~23% and ~73%, respectively, and the energy consumption will be decreased by ~63% for low-resistance state restoration, compared with the prior art initialization-and-overwrite-7T1R at nMOS typical corner and pMOS typical corner in Taiwan Semiconductor Manufacturing Company's 65-nm technology. In addition, with the column-shared structure, the area penalty is cheerfully acceptable.
Chunyu Peng, Songsong Xiao, Wenjuan Lu, Xiulong Wu, Junning Chen, Zhi-Ting Lin
IEEE Trans. Very Large Scale Integr. Syst.5
2016 A yield-enhanced global optimization methodology for analog circuit based on extreme value theory
Minghua Li, Guanming Huang, Xiulong Wu, Liuxi Qian, Xuan Zeng 0001, Dian Zhou
Sci. China Inf. Sci.3
2015 Analyzing and modeling mobility for infrastructure-less communication
Zhi-Ting Lin, Xiulong Wu
J. Netw. Comput. Appl.2
2012 Bitline Leakage Current Compensation Circuit for High-Performance SRAM Design
abstract
The leakage current existing in the bitline of SRAM has attracted more and more concerns for the operation of high-performance SRAM design, especially with the decrease of the threshold voltage of the transistor for high-performance demand, the leakage current would increase exponentially. The increased leakage current may slowdown the performance of the read operation of SRAM because the existence of the leakage current in the bitline may postpone the time to resolve the sufficient differential bitline voltage for SA to sense correctly. In this paper, a new bitline leakage current compensation circuit has been proposed. Different from the traditional technique, the proposed bitline leakage current compensation circuit cancels the "pre-determined" leakage current compensation process. Therefore, it dodges the dilemma where the performance of the SRAM may be degraded in some circumstances if the bit-line leakage current is compensated pre-determinedly. The simulation results show that by adopting the proposed compensation circuit, the time needed to develop 1/2 VDD can be reduced by almost 126.4% under the tt process corner.
Ruixing Lu, Na Bai, Baitao Lv, Jiafeng Zhu, Xiulong Wu
NAS5