Zhi-Ting Lin

dblp:81/2960 · also Zhiting Lin · DBLP profile ↗
← Back
47ranked-venue papers
17as first author
33since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 38 · 10 first-author · 33 since 2021Computer networks · 6 · 5 first-authorDatabases, data management, data science and information retrieval · 2 · 2 first-authorArtificial intelligence and machine learning · 1
YearPublicationVenuePosition
2026 A 2RW Dual-Port 8T-SRAM Macro with Bitline Leakage Current Tracking and Read-Write Arbitration
Chenghu Dai, Junbo Chen, Zaihang Zhang, Licai Hao, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu
ISCAS9
2026 A Multiplication-Free Floating-Point CIM Architecture with a 5-Bit Approximate Squaring Circuit
Xin Li 0099, Juntao Ge, Miao Long, Fugui Jiang, Chenglong Duan, Yang Yang 0025, Yu Liu 0113, Zhi-Ting Lin
ISCAS10
2026 A Floating-Point CIM Macro Featuring Asymmetric Exponent Encoding and Adaptive Mantissa Truncation for High-Efficiency AI Edge Computing
Zhi-Ting Lin, Rongtao Li, Yu Liu 0113, Xin Li 0099, Xiulong Wu
ISCAS1
2026 An offset-compensation capacitor-coupled DRAM sense amplifier with symmetric sensing and high PVT stability
Chenghu Dai, Jing Lv, Yongqi Qin, Chunyu Peng, Xin Li 0099, Yu Liu 0113, Xiulong Wu, Zhi-Ting Lin
Integr.9
2026 Analysis and Design of Memory Testing Algorithm for Computing-in-Memory Using MBIST
abstract
Computing-in-memory (CIM), as a novel computing architecture for the future, effectively overcomes the bottlenecks in the von Neumann architecture. The CIM architecture embeds logic into the memory array to reduce the data transfer between the processor and memory. However, embedding logic into the memory array increases the test complexity. In this study, we offer a comprehensive examination of the challenges associated with CIM and introduce a novel March-like test algorithm, named March CC, tailored for CIM chips. Computational elements are added to the read/write operation sequences, combining the tests in memory mode and computing mode into one step, which significantly improves the test efficiency. In comparison to the traditional March C− test algorithm, the proposed March CC test algorithm, with a complexity of only 10 N , enhances the fault coverage from 66.7% to 79.8% for six common single-cell fault (SCF) models and nine common double-cell fault (DCF) models. Furthermore, the March CC algorithm demonstrates good compatibility and is applicable to various memory configurations, such as SRAM, RRAM, and MRAM CIM architectures.
Zhi-Ting Lin, Siyan Li, Qiushi Feng, Changxin Yue, Yuanyang Wang, Yunlong Liu 0006, Yu Liu 0113, Licai Hao, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Xiulong Wu
ACM J. Emerg. Technol. Comput. Syst.1
2026 Time-Domain SRAM-CIM Macro With Dual-Edge Temporal Fused Accumulation for Signed 8-bit Precision MAC
Wenjuan Lu, Xiaobo Gong, Kang Meng, Xiaohang Chen, Jiating Guo, Lijun Guan, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu, Chunyu Peng
IEEE Trans. Circuits Syst. I Regul. Pap.9
2026 A T8T-SRAM Computing-in-Memory Macro for Ternary Deep Neural Networks and Boolean Logic Computations
abstract
Deep neural networks (DNNs) play important roles in artificial intelligence applications and show hungry computility and power demands. Compared with binary neural networks (BNNs), ternary neural networks (TNNs) have higher representation and adaptive abilities and balance the inference accuracy and computing efficiency between DNNs and BNNs. This article proposed a T8T-SRAM computing-in-memory (CIM) macro to achieve Boolean logic operations and MAC operation of ternary activation and ternary weight. The proposed T8T-SRAM bitcell has a separate read and write path, and can avoid the read disturb issue. In Boolean logic operation mode, the T8T-SRAM macro can achievenand,nor,xnor, andxoroperations with redundant rows, reducing the additional reference voltage generation circuit. In the MAC mode, the result is quantized by an embedded column analog-to-digital converter (ADC), which uses activation refresh to reduce weight changing. In 28-nm CMOS technology, under 0.5-V array supply voltage and 0.9-V peripheral supply voltage, simulation results manifest that the MAC results have good linearity, and feasibility of Boolean logic operation. The proposed T8T-SRAM macro realizes MAC operation of 16 ternary activations and 16 ternary weights with 333.99–816.1-TOPS/W energy efficiency and 61.9-TOPS/mm2area efficiency. Using an ResNet-18 network for the inference of MNIST, and CIFAR-10 datasets, the accuracies were 99.06% and 85.76% with a ternary activation and ternary weight.
Chenghu Dai, Zihua Ren, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.9
2026 A 28 nm 1.3 TFLOPS/mm2 Floating-Point SRAM-Based CIM Macro With Asynchronous Normalization and Parallel Sorting Alignment for AI-Edge Chip
abstract
State-of-the-art AI edge devices require floating-point (FP) multiply-accumulate (MAC) operations with high-energy efficiency and inference accuracy. FP computing in-memory (FP-CIM) has a broader range of applications compared to integer CIM. However, FP-CIM can incur greater power, delay, and area overheads than integer CIM due to the inherent complexity of FP computational flow. In this article, we introduce a new method for asynchronous exponent normalization and parallel mantissa alignment. This approach allows us to add exponents and find the maximum sum simultaneously. We also replace the traditional subtraction and shifting for mantissa alignment with a cross-structure maximum-finding method, enabling FP-CIM to be achieved with lower delay, area, and power overheads. The macro is designed in TSMC 28 nm process, with a memory size of 6 Kb, a layout area of 0.067 mm2, and an area efficiency of 1.3 TFLOPS/mm2. Simulation results show that the macro computational frequency and energy efficiency can reach 150 MHz and 12.8 TFLOPS/W, respectively, at 900 mV, while performing FP- MAC operations.
Zhi-Ting Lin, Miao Long, Yang Yang 0025, Lintao Chen, Yu Liu 0113, Xin Li 0099, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.1
2026 A Digital FP CIM Macro With Cascaded Row-Elimination for Exponent Comparison and In-Memory Mantissa Sparsity Detection
abstract
Floating-point (FP) computing-in-memory (CIM) addresses the energy efficiency bottleneck of von Neumann architectures and fixed-point CIM in high-accuracy neural network training/inference. However, existing FP CIM designs still suffer from limited exponent-path parallelism and low mantissa-path energy efficiency. This article proposes a novel FP CIM architecture using: 1) a cascaded row-elimination comparison mechanism for single-cycle, high-parallelism exponent max-value comparison; 2) an in-memory sparse feature detection method that skips redundant mantissa computing modules based on different input/weight mantissa patterns to reduce energy consumption; and 3) separated exponent/mantissa CIM modules enabling pipelining and a mantissa bit-extension strategy optimizing accuracy, area, and energy efficiency. Simulation results of a 28-nm 19-kb macro show a 496-MHz operating frequency at 0.9 V and a peak energy efficiency of 12.88 TFLOPS/W.
Wenjuan Lu, Kang Meng, Xiaobo Gong, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.9
2026 A 28-nm 9-kb SRAM Computing-in-Memory Macro With Segmented Charge Sharing for Multimode MAC Operations
abstract
Computing-in-memory (CIM) is a novel approach to solve the von Neumann bottleneck and improve energy efficiency and throughput. This article presents an SRAM-CIM macro based on segmented charge sharing to support multimode multiply-and-accumulate (MAC) operations, including binary weight network (BWN) MAC, ternary weight network (TWN) MAC, and multibit MAC operations. Signed 9-bit input MAC operations can be supported in BWN and TWN networks through multiple cycles. The multibit MAC operations are realized by the cooperation of two cell structures, which eliminates the need to convert negative numbers into complements and reduces the area and power consumption of the chip. Additionally, the proposed segmented charge-sharing scheme increases the speed of accumulation and improves the overall energy efficiency of the chip. The proposed 9-kb macro is implemented in 28-nm CMOS technology with an energy efficiency of 30.74 TOPS/W and an area efficiency of 1.15 TOPS/mm2in multibit MAC mode. At the system level, a ResNet-based model achieves an accuracy of 94.59% on the CIFAR-10 dataset and 74.64% on the CIFAR-100 dataset, demonstrating the effectiveness of the proposed CIM architecture for practical neural network inference. Compared to prior designs, it not only supports additional computing modes but also demonstrates notable improvements in both energy and area efficiency.
Wenjuan Lu, Lening Tan, Tianchen Xue, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.6
2026 A 28-nm 16-Kb SRAM Computing-in-Memory Macro With Dual Bitline Computing for Boolean Logic Operation and BWN MAC Operation
abstract
By integrating computation directly within memory structures, computing-in-memory (CIM) effectively mitigates the von Neumann bottleneck, offering significant improvements in energy efficiency and system throughput. This letter presents a 16-kb static random access memory (SRAM) CIM macro in 28-nm complementary metal–oxide–semiconductor (CMOS), featuring a novel dual bitline computing cell (DBCC) architecture. The design supports three operational modes: conventional memory access, Boolean logic operations, and binary weight network (BWN) multiply-and-accumulate (MAC) computations. The DBCC architecture addresses key critical limitations of prior works by: 1) storing both positive and negative weights within a single 6T-SRAM cell, improving array utilization by$2\times $compared to segregated storage schemes; 2) achieving consistent discharge rates for both polarities through symmetric pull-down paths, enhancing calculation accuracy [INL <3 least significant bit (LSB) across corners] and process variation tolerance; and 3) enabling differential voltage-based Boolean logic operations without external reference voltage generation, thereby reducing peripheral circuitry. The design features a compact 2:1 multiplexer (2:1 MUX)-based pulsewidth modulator (PWM) for 5-b signed input encoding and demonstrates 55.8–204.8-TOPS/W energy efficiency in 28-nm CMOS. Experimental results on a 16-kb array show stable operation across PVT variations while supporting both memory functions and in situ computation for BWNs. System-level evaluation demonstrates its practicality, achieving competitive inference accuracies of 97.59% on MNIST and 89.08% on CIFAR-10.
Wenjuan Lu, Lubin Xiang, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.6
2026 A 9T SRAM Computation-in-Memory Architecture With High-Precision MAC, Enhanced Bitline Voltage Margin, and Improved Frequency Performance Over Conventional Architectures
abstract
To address the data-intensive demands of modern artificial intelligence (AI) systems, computation-in-memory (CIM) based on static random-access memory (SRAM) has emerged as a promising solution by integrating computing functionality within memory arrays. However, conventional SRAM CIM architectures face two key limitations: low output resistance in single-transistor transmission paths and voltage instability on charge-sharing bitlines. These limitations collectively degrade computational accuracy to 4–5 LSB-level integral nonlinearity (INL), restricting practical deployment. This work proposes a regulated-cascode 9T SRAM cell that enhances analog computation accuracy using a high-impedance transmission path through a cascode configuration and stabilizing the discharge amount of the bitline from a single cell via active feedback regulation. Implemented in Semiconductor Manufacturing International Corporation (SMIC) 55-nm CMOS technology, the proposed cell demonstrates 1.31LSB INL at 400-mV bitline swing (68.4% improvement versus 4–5 LSB baselines), achieving 66.7% voltage utilization efficiency compared with the conventional 50% limit and 23.04% frequency improvement is achieved compared with the conventional architecture. It also achieves an energy efficiency of 18.47 fJ/bit and a compact area of$2.655\times 1.175~\mu $m, while demonstrating a classification accuracy of 97.7% on the MNIST dataset.
Hongbiao Wu, Zhi-Ting Lin
IEEE Trans. Very Large Scale Integr. Syst.2
2026 A CIM Macro Embedded With Sign Operations for Parallel Signed Multibit Multiplication-and-Accumulation Using Hybrid Cell Array
Jin Zhang 0036, Zhongzhen Tong, Qiang Zhao 0007, Chunyu Peng, Wenjuan Lu, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.8
2025 A Floating-Point SRAM-based CIM Macro with Asynchronous Normalization and Parallel Sorting Alignment
abstract
Floating-point computing-in-memory (FP-CIM) has a broader range of applications compared to integer CIM. However, FP-CIM can incur greater power, delay, and area overheads than integer CIM due to a more complex computational flow. In this paper, we introduce a new method for asynchronous exponent normalization and parallel mantissa alignment. This approach allows us to add exponents and find the maximum sum simultaneously. We also replace the traditional subtraction and shifting for mantissa alignment with a time-cycle lookup method, enabling FP-CIM to be achieved with lower delay, area, and power overheads. The macro is designed in the TSMC 28nm process, with a memory size of 6Kb, a layout area of 0.067mm2, and an area efficiency of 1.3TFLOPS/mm2. Simulation results show that the macro computational frequency and energy efficiency can reach 150MHz and 12.8TFLOPS/W, respectively at 900mV.
Zhi-Ting Lin, Dongcheng Wang, Rongtao Li, Shichen Yu, Yu Liu 0113, Xin Li 0099, Xiulong Wu
ISCAS1
2025 TSCIM: A 28nm Transposed Stochastic CIM Macro for On-Chip Training and Inference
abstract
This work introduces a novel Transposed Stochastic Computing-in-Memory (TSCIM) macro designed to enhance the efficiency of on-chip training and inference. The macro incorporates a novel stochastic quantization strategy and utilizes a transposed separated wordline SRAM to enable multi-bit signed MAC operations. Furthermore, a stochastic adder tree is utilized to minimize area and power consumption overhead. The design includes a 4Kb SRAM CIM macro implemented in 28 nm CMOS technology. Simulation results show that the power consumption of the stochastic accumulation circuit (SAC) is reduced by 63.6%, while the area overhead is decreased by a factor of 7.73 compared to designs using full adder (FA) adder trees. Additionally, the computation latency is decreased by 16× compared to traditional stochastic circuits. The TSCIM macro can achieve a peak energy efficiency of 63.02 TOPS/W and an area efficiency of 15.54 TOPS/mm2.
Yu Liu 0113, Yang Lou, Kangkang Mao, Xin Li 0099, Chenghu Dai, Xiulong Wu, Zhi-Ting Lin
ISCAS7
2025 MTJ based temperature compensated beta multiplier Voltage Reference
abstract
This article mainly explores the interaction between CMOS devices and Magnetic Tunnel Junction (MTJ) devices in terms of temperature characteristics, aiming to achieve a CMOS beta multiplier circuit that combines low power consumption and wide temperature adaptability, making it a stable reference voltage source. The proposed design utilizes the Tunneling Magneto Resistance (TMR) effect of MTJ to compensate for the performance mismatch caused by temperature changes in CMOS. This design adopts TSMC 28nm CMOS craft, which can generate a reference voltage with a linearity of 0.57%/V, a temperature coefficient of 43.6ppm/°C, and a stable voltage of 441.6mV at a minimum supply voltage of 0.6V and a temperature range of 10~110 °C. The noise of this voltage at a frequency of 10Hz is 69.7uV/sqrt (Hz), the power suppression ratio is -62.1dB, and the power consumption is 4.092nW.
Yongliang Zhou, Yingxue Sun, Jingxue Zhong, Chengxing Dai, Weizhe Tan, Chunyu Peng, Xin Li 0099, Zhi-Ting Lin, Xiulong Wu
ISCAS9
2025 MTJ based Temperature-Adaptive VCO (TAVCO) for Compensating CP-PLL Frequency Drift
abstract
The Charge Pump Phase-Locked Loop (CP-PLL) is a commonly utilized component in contemporary mixed-signal electronic systems. It is widely employed for clock generation, synchronization, and frequency synthesis in both digital and wireless functionalities. However, the frequency accuracy of oscillators can be adversely affected by variations in frequency across a broad temperature range. To address this issue, the voltage-controlled oscillator designed in this study employs a four-stage differential delay structure, chosen for its simple circuit architecture, favorable control linearity, and low noise characteristics. This research integrates the temperature behaviors of Complementary Metal-Oxide-Semiconductor (CMOS) and Magnetic Tunnel Junction (MTJ) technologies, utilizing 28nm CMOS technology to enhance the frequency stability of ring oscillators effectively. Simulation results indicate that frequency drift is reduced by 92% within the temperature range of -80°C to 125°C.
Yongliang Zhou, Jingxue Zhong, Yingxue Sun, Chengxing Dai, Weizhe Tan, Chunyu Peng, Wenjuan Lu, Xin Li 0099, Zhi-Ting Lin, Xiulong Wu
ISCAS9
2025 A Low-Cost and Triple-Node-Upset Self-Recoverable Latch Design With Low Soft Error Rate
abstract
With the decrease in feature size of transistors, latches are more sensitive to single-event multiple node upset (MNU), including double node upset (DNU) and triple node upset (TNU). However, the reported TNU self-recoverable (TNUR) latches are facing problems with large areas and power consumption. Based on the polarity design, this article proposes a low-cost TNUR latch (LCTRL) with a low soft error rate (SER) in 28-nm CMOS technology. The proposed LCTRL mainly consists of four interlocked modules and a clock-gated inverter. Compared with the state-of-the-art TNUR latches, including LCTNURL, IHTRL, FATNU, and TRLW, the power consumption, D-Q delay, CLK-to-Q delay, area, and the power-delay–area product (PDAP) of the proposed LCTRL are reduced by 55.09%, 38.64%, 42.93%, 44.65%, and 83.50%, respectively. Due to the polarity design, the SER of the proposed LCTRL is the smallest among compared latches, which suggests that the proposed LCTRL is suitable for use in radiation environments.
Licai Hao, Lang Tian, Hao Wang 0239, Shiyu Zhao 0004, Qiang Zhao 0007, Chunyu Peng, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.8
2025 Full-Array Boolean Logic CIM Macro With Self-Recycling 10T-SRAM Cell for AES Systems
abstract
Computing in memory (CIM), which alleviates the need to transfer a large amount of data between processor and memory, significantly reducing latency and energy consumption, is a promising new computing architecture for addressing the von Neumann bottleneck problem. This article proposes a CIM array structure composed of self-recycling 10T static random access memory (SRAM) cells, which can realize orthogonal data writing, and multiple Boolean logical operations for the entire array. The self-recycling and full-array activation characteristics are extremely suitable for accelerating diverse data processing algorithms such as the Advanced Encryption Standard (AES). A 4-kb SRAM is implemented in 55-nm CMOS technology to verify the effectiveness of the design. Compared with other state-of-the-art architectures, the throughput and the operating frequency of the proposed CIM macro are increased to 843 GOPS/kb ($2.64\times $) and 823.7 MHz ($2.6\times $), respectively. The energy efficiency reaches 246.9 TOPS/W. When applied to the AES, the energy consumption is 35.77% less than the digital CIM architecture that is not self-recycling.
Xin Li 0099, Lintao Chen, Yang Lou, Baofa Wu, Jiajun Long, Yongliang Zhou, Chunyu Peng, Xiulong Wu, Zhi-Ting Lin
IEEE Trans. Very Large Scale Integr. Syst.11
2025 A 28-nm 9T1C SRAM-Based CIM Macro With Hierarchical Capacitance Weighting and Two-Step Capacitive Comparison ADCs for CNNs
abstract
In the realm of charge-domain computing-in-memory (CIM) macros, reducing the area of capacitor ladder and analog-to-digital converter (ADC) while maintaining high throughput remains a significant challenge. This brief introduces an adjustable-weight CIM macro designed to enhance both energy efficiency and area efficiency for convolutional neural networks (CNNs). The proposed architecture uses: 1) a customized 9T1C bit cell for sensing margin improvement and bidirectional decoupled read ports; 2) a hierarchical capacitance weighting (HCW) structure that achieves a weight accumulation of 1/2/4 bits with less capacitance area and weighting time; and 3) a two-step capacitive comparison ADCs (TC-ADCs) readout scheme to improve area efficiency and throughput. The proposed 8-kb static random address memory (SRAM) CIM macro is implemented using 28-nm CMOS technology. It can achieve an energy efficiency of 224.4 TOPS/W and an area efficiency of 21.894 TOPS/mm2, and the accuracies on MNIST, CIFAR-10, and CIFAR-100 datasets are 99.67%, 89.13%, and 67.58% with a 4-b input and 4-b weight.
Zhi-Ting Lin, Runru Yu, Miao Long, Yu Liu 0113, Jianxing Zhou, Qingchuan Zhu, Yue Zhao 0029, Lintao Chen, Chunyu Peng, Qiang Zhao 0007, Xin Li 0099, Chenghu Dai, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.1
2025 High-Reliability and High-Throughput CIM 10T-SRAM for Multiplication and Accumulation Operations With 274.3 GOPS and 200-237.5 TOPS/W
abstract
Artificial intelligence (AI) is extensively applied in natural language processing, image matching, and image recognition, with convolutional neural networks (CNNs) being crucial. Computing-in-memory (CIM) utilizing static random access memory (SRAM) can enhance the CNN performance. However, this faces issues such as multibit signed data processing, read corruption of traditional SRAM arrays, and increased area overhead due to increased capacitor weighting. This article proposes a 10T-SRAM macro tailored for CNN multiply-accumulate calculation (MAC) computation in image processing. It enables high-throughput full-array operations, with added dual ports facilitating input of multibit data with signed bits. The 10T-SRAM cell features a read-write separation channel, mitigating read disturbance issues seen in dual-port 8T-SRAM arrays or 6T-SRAM arrays. Incorporating redundant columns in the array for charge sharing and weighting conserves area and boosts circuit reliability. In the 28-nm CMOS simulation environment, the proposed architecture achieves a throughput of 274.3 GOPS and an energy efficiency of 200–237.5 TOPS/W, surpassing literature-reported figures by several times.
Wenjuan Lu, Lubin Xiang, Chunyu Peng, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.6
2025 A 28 nm Dual-Mode SRAM-CIM Macro With Local Computing Cell for CNNs and Grayscale Edge Detection
abstract
With the rise of artificial intelligence (AI), neural network applications are growing in demand for efficient data transmission. The traditional von Neumann architecture can no longer keep pace with modern technological needs. Computing-in-memory (CIM) is proposed as a promising solution to address this bottleneck. This work introduces a local computing cell (LCC) scheme based on compact 6T-SRAM cells. The proposed circuit aims to enhance energy efficiency and reduce power consumption by reusing the LCC. The LCC circuit can perform the multiplication of a 2-bit input with a 1-bit weight, which can be applied to convolutional neural networks (CNNs) with the multiply-accumulate (MAC) operations. Through circuit reuse, it can also be used for multibit multiply operations, performing 2-bit input multiplication and 1-bit weight addition, which can be applied to grayscale edge detection in images. The energy efficiency of the SRAM-CIM macro achieves an energy efficiency of 46.3 TOPS/W under MAC operations with input precision of 8-bits and weight precision of 8-bits, and up to 389.1–529.1 TOPS/W under the calculation in one subarray with an input precision of 2-bits and a weight precision of 1-bit. The estimated inference accuracy on CIFAR-10 datasets is 90.21%.
Chunyu Peng, Xiaohang Chen, Mengya Gao, Jiating Guo, Lijun Guan, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.7
2025 A 28-nm Cascode Current Mirror-Based Inconsistency-Free Charging-and-Discharging SRAM-CIM Macro for High-Efficient Convolutional Neural Networks
abstract
Computing-in-memory (CIM) is an emerging approach to alleviate the von Neumann bottleneck and enhance energy efficiency and throughput. This brief introduces a 16-Kb static random access memory (SRAM) CIM macro for convolutional neural networks (CNNs), featuring a cascode current mirror-based inconsistency-free computing circuits (CICCs). The bias voltage of CICC is provided by a cascode current mirror (CCM) circuit. The proposed architecture improves the consistency and linearity of bitline (BL) charge and discharge rates in the analog current domain, enhancing computational accuracy. Additionally, the charge and discharge on the BLs represent the positive or negative calculation result, eliminating the need for extra encoding and logic circuits to handle sign bits. The SRAM-CIM macro achieves an energy efficiency of 59.1–134.0 TOPS/W and a throughput of 0.41 TOPS in a 28-nm CMOS technology, and the estimated inference accuracy on MNIST and CIFAR-10 datasets is 96.5% and 91.4%, respectively, with 5-bit input precision and 1-bit weight precision.
Chunyu Peng, Jiating Guo, Shengyuan Yan, Xiaohang Chen, Wenjuan Lu, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.8
2025 A High-Performance and High-Robustness Triple-Node-Upset Tolerant Latch Based on Redundant-Node Hardening
abstract
In response to the issues of high cost, large overhead, and limited node fault tolerance in current latch hardening techniques, this article proposes a latch circuit resistant to triple-node-upset (TNU) based on redundant-node hardening technology. This latch comprises eight 1P2N modules interlocked, with its output isolated by two levels of C-elements (CEs), achieving tolerance to TNU. The performance of the redundant-node reinforcement TNU tolerant latch (RNRTTL) was simulated and verified using CMOS 65 nm technology. The simulation results indicate that the RNRTTL circuit has a D-Q delay of 14.14 ps, static power consumption of$4.03~\mu $w, an area of$32.87~\mu $m2, and an area-static power-D–Q delay-product (APDP) of 1873, respectively. Compared to the triple-node upset tolerant latches TTLL, TNU-latch, TNURL, and HLTNURL reported in the current literature, the proposed latch demonstrates an average reduction of 219.9%, 164.9%, 150.7%, and 2464.8% in D-Q delay, static power consumption, area, and APDP, respectively, indicating that the RNRTTL latch has superior comprehensive performance; furthermore, a series of 2000 Monte Carlo (MC) simulations on the node group$\langle $Q, X0, X$8\rangle $reveal that the proposed latch circuit possesses good stability, making it suitable for harsh radiation environments.
Qiang Zhao 0007, Qingyi Liu, Licai Hao, Xin Li 0099, Shengyue Zhang, Chunyu Peng, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.8
2024 SRAM-Based Digital CIM Macro for Linear Interpolation and MAC
abstract
Linear interpolation is widely used in algorithms such as image segmentation, but the existing compute-in-memory (CIM) architectures cannot satisfy the needs of linear interpolation. This paper proposes a CIM macro based on static random-access memory (SRAM) that implements linear interpolation for the first time. Swing multiplication and accumulation are proposed in this paper for linear interpolation operations. In addition, the proposed circuit can be used for multiply-and-accumulate (MAC) operation and supports parallel updating and computing. A new adder tree by reused the adders inside the multiplier to implement the accumulation operation. The proposed circuit can improve the area efficiency as no additional adder tree circuit is required. The design is implemented in a 28 nm process and area efficiency can achieve 0.059 TOPS/mm2for MAC operation. When operating with an 8-bit linear interpolation, this CIM macro achieves access times of 6−9 ns and energy efficiencies of 11.13−17.72 TOPS/W. Under MAC operation with 8-bit inputs and weights, this CIM macro achieves access time of 6.36−9.47 ns and energy efficiency of 4.61−8.36 TOPS/W for 19-bit outputs.
Zhi-Ting Lin, Yunlong Liu 0006, Yaling Wang, Yue Zhao 0029, Chunyu Peng, Xiulong Wu
ISCAS1
2024 A High Throughput In-MRAM-Computing Scheme Using Hybrid p-SOT-MTJ/GAA-CNTFET
abstract
Silicon-based semiconductor transistors are approaching their physical limits due to shrinking feature sizes. Simultaneously, traditional silicon-based von Neumann architectures exhibit significant latency and power consumption issues in data-centric applications, such as the Internet of Things and artificial intelligence. To tackle these challenges, this study introduces a novel approach: Magnetoresistance Random Access Memory (MRAM) computing in-memory (CIM) using gate-all-around carbon nanotube field-effect transistors (GAA-CNTFET). The proposed MRAM array comprised three transistors and one perpendicular magnetic anisotropy spin-orbit torque magnetic tunnel junction (p-SOT-MTJ) (3T1M) cell and achieves full-array Boolean logic operations and half/full-adder operations. The calculated results can be stored in-situ during the computing phase without requiring additional peripheral circuits. A 16 Kb MRAM was simulated in both GAA-CNTFET/p-SOT-MTJ and 14-nm FinFET/p-SOT-MTJ technologies to examine the effectiveness of the proposed design. Compared to its 14-nm FinFET/p-SOT-MTJ counterparts, the write and computing latencies of the GAA-CNTFET/p-SOT-MTJ CIM macro were reduced by approximately 21% and 20.6%, respectively, while the read and computing energy consumption by approximately 45.3% and 24.7%, respectively. Moreover, the proposed in-memory Boolean logic throughput was 8192 GOPS, which was approximately 160–250 times higher than that of existing CIM solutions, in which only two rows of word lines can be activated.
Zhongzhen Tong, Yunlong Liu 0006, Xinrui Duan, Suteng Zhao, Chenghang Li, Zhi-Ting Lin, Xiulong Wu, Zhaohao Wang, Xiaoyang Lin
IEEE Trans. Circuits Syst. I Regul. Pap.8
2024 A Computing In-Memory Multibit Multiplication Based on Decoupling and In-Array Storing
abstract
Multiplications are basic operations of neural networks. Therefore, multiplication results are crucial in analyzing the operating process of neural networks. However, the multiplication strategies are generally based on analog-domain circuits, and the results are in a multiply-and-accumulate (MAC) form. The result of each multiplication in MAC cannot be distinguished accurately using these strategies. Therefore, we proposed an in-memory multibit multiplication based on the decoupling and in-array storage strategy to overcome this problem, and the core module is the 10T1C SRAM cell. Multibit multiplications are decoupled by a series of logical operations. Therefore, in the analysis mode, multiplication results can be saved and outputted in the normal read mode without requiring additional storage. When executing the neural network, the operation results are stored in the cells. Hence, the operands stored in the array are retained. Accumulation operations are completed based on the charge-sharing technology; thus, the linearity of accumulation is high. We simulated and analyzed the performance of the proposed circuit in a 28 nm CMOS process. The absolute value of integral nonlinearity is at most 0.29. Further, due to high data operation parallelism, the throughputs of the logical operation and MAC are up to 6307.8 and 802.8 GOPS, respectively.
Jin Zhang 0036, Zhongzhen Tong, Hao Wang 0239, Qiang Zhao 0007, Jiaqun Wang, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Circuits Syst. I Regul. Pap.8
2024 Low-Cost and Highly Robust Quadruple Node Upset Tolerant Latch Design
abstract
This article proposes an exceptionally reliable and low-cost quadruple node upset tolerant latch ($LC$-QNUTL) suitable for the 65 nm CMOS technology. The innovative$LC$-QNUTL latch is primarily composed of three soft-error-immune (SEI) static random-access memory (SRAM) cells and a triple-level C-element (CE) unit, which includes five two-input CE and a clock-gating (CG)-based two-input CE. The SEI SRAM cell utilizes polarity hardening technology and source-isolation technology, significantly reducing the number of sensitive nodes and enhancing the latch’s stability. By using the high-speed transmission gate (TG) technology and stacked structures, the proposed latch offers minimal overhead in terms of delay and power consumption, yielding an improved power delay area product (PDAP). When compared to contemporary quadruple node upset (QNU)-tolerant latch designs (including HLMR, 4NUHL, and LDAVPM), the new design offers substantial improvements—29.53% less delay, 80.09% reduced power consumption, 58.52% smaller silicon area, and 433.43% improved comprehensive PDAP on average. Furthermore, simulation results demonstrate that the$LC$-QNUTL latch exhibits reduced sensitivity to process, voltage, and temperature (PVT) variations, thus providing superior reliability, which makes it an ideal choice for safety-critical applications.
Licai Hao, Yaling Wang, Yunlong Liu 0006, Shiyu Zhao 0004, Wenjuan Lu, Chunyu Peng, Qiang Zhao 0007, Yongliang Zhou, Chenghu Dai, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.12
2024 Soft-Error-Immune Quadruple-Node-Upset Tolerant Latch Based on Polarity Design and Source-Isolation Technologies
abstract
A soft-error-immune quadruple-node-upset tolerant latch (SEI-QNUTL) with a low delay and high performance is proposed using 65-nm CMOS technology. The proposed SEI-QNUTL design consists of three soft-error-immune static random access memory (SEI-SRAM) cells. Furthermore, each SEI-SRAM cell employs polarity design and source-isolation technology to reduce the number of sensitive nodes and enhance the reliability of the latch. Compared with state-of-the-art quadruple-node-upset (QNU) tolerant latches [including high-performance and low-cost single-event multiple-node-upsets resilient (HLMR), QNU tolerant latch (QNUTL), and Latch Design and Algorithm-based Verification Protected against Multiple-Node-Upsets (LDAVPM)], the proposed SEI-QNUTL design reduces (on average) the area, delay, and area-power-delay-product (APDP) by 47.0%, 25.0%, 46.5%, and 66.3%, respectively. Extensive variation analysis validates that the SEI-QNUTL design is less sensitive to process, voltage, and temperature (PVT) variations regarding power consumption and delay. Furthermore, Monte Carlo (MC) simulations show that the proposed latch exhibits high reliability when performing data storage. Compared with the existing latches, the SEI-QNUTL design makes a good tradeoff among delay, power, and area, and it can thus be used in safety-critical applications.
Licai Hao, Chenghu Dai, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng, Yongliang Zhou, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.8
2023 High Restore Yield NVSRAM Structures With Dual Complementary RRAM Devices for High-Speed Applications
abstract
Static random access memory (SRAM) plays a key role in the overall performance of electronic systems because of its rapid data processing and transmission speed; however, when the system power supply is cut off, the data stored in the nodes are lost. Thus, this article proposes four nonvolatile SRAM (NVSRAM) cells that use unilateral or bilateral structures with dual complementary series resistive random access memory (RRAM) devices. It is found that the read, write, and hold static noise margins (HSNMs) are comparable with those of the standard 6T-SRAM. Moreover, the store and restore operations operate in parallel at high speed. The store operation delay is only 6 ns for unilateral structures and 5 ns for bilateral structures, and the restore delay is only 10 ns for unilateral structures and 6 ns for bilateral structures. The maximum power consumption among the four structures for storing and restoring a “1” are 1.545 pJ/bit and 134.5 fJ/bit, respectively. Furthermore, the dual complementary series resistor structures can achieve a high restore yield at a resistance ratio of 1.5. Therefore, a high restore yield can be achieved even with large resistance fluctuations caused by the voltage, time, and process.
Zhi-Ting Lin, Xiulong Wu, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng
IEEE Trans. Very Large Scale Integr. Syst.1
2023 A Fully Digital SRAM-Based Four-Layer In-Memory Computing Unit Achieving Multiplication Operations and Results Store
abstract
The separation of memory and arithmetic logic unit (ALU) in the von Neumann computing architecture hinders the development of big data and high-performance computing. In-memory computing (IMC) as a new computation method significantly reduces the latency and power consumption of data processing. In this study, we propose a fully digital static random access memory (SRAM)-based IMC architecture, which has the following advantages: 1) it simplifies multiplication to multicycle addition operations, reuses logic cells, and reduces hardware overhead; 2) by adding a pair of nMOS transistors to achieve internal write-back, the computational efficiency is improved, and at the same time, the final result of the multiplication can be stored locally, eliminating the need to read the computational result immediately; and 3) this scheme can be easily expanded to multiplication operations with different bit widths, which provides good scalability. A 4-kb SRAM-IMC macro chip is manufactured using the SMIC 55-nm technology to realize 4-bit multiplication, with an energy efficiency of 51.4 TOPS/W (0.9 V) and a throughput of 234.3 GOPS/mm2. The proposed multiplication–accumulation architecture is applied to a neural network, which achieves 98.7% accuracy with the Mixed National Institute of Standards and Technology database (MNIST) dataset.
Zhi-Ting Lin, Shaoying Zhang, Jianping Xia, Yunwei Liu, Kefeng Yu, Zhongzhen Tong, Xiulong Wu, Wenjuan Lu, Chunyu Peng, Qiang Zhao 0007
IEEE Trans. Very Large Scale Integr. Syst.1
2023 In-Memory Transposable Multibit Multiplication Based on Diagonal Symmetry Weight Block
abstract
A possible approach to overcome the von Neumann bottleneck and meet the increasing demand for better computing performance is to computing in-memory (CIM). The results of the in-memory calculations are primarily reflected in the vertical bitline (BL) analog voltage. However, the nonlinearity of the BL discharge deteriorates with the increase in discharge voltage. In this study, we propose a diagonal symmetry weight block (DSWB) based on an eight-transistor (8T) static random access memory (SRAM) that can achieve multibit transposable operations. In addition, to guarantee linearity and complete multibit multiplication operations, we propose a cascode current mirror (CCM)-based multiplier. To achieve low-overhead and more efficient quantification, our proposed CIM macro uses a counter-type quantization circuit to read out the analog calculation results. We simulated the performance of the proposed 8T SRAM in a 28-nm complementary metal–oxide–semiconductor process. The integral nonlinearity (INL) of the proposed CCM-based CIM decreased by approximately 54.4% compared with the traditional CIM. Furthermore, the proposed in-memory multibit multiplication throughput density was 6.74 GOPS/kb; this throughput density improvement is approximately 3.3–10.5 times higher than the existing CIM works.
Zhongzhen Tong, Yue Zhao 0029, Jin Zhang 0036, Zhi-Ting Lin, Xiaoyang Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.4
2022 Configurable Memory With a Multilevel Shared Structure Enabling In-Memory Computing
abstract
Frequent to-and-from data transfers in the von Neumann architecture limit the overall throughput. One of the promising approaches used to overcome von Neumann bottleneck is in-memory computing (IMC) that aims to embed computing in memory to reduce the transfer of memory-processor data. This study proposes a configurable 6-transistor (6T) static random access memory (SRAM) array with a multilevel shared structure for IMC. A multilevel shared structure can effectively improve the utilization rate of the module. In addition to the conventional SRAM operation, the configurable structure can also perform the sum of absolute differences (SAD) and Hamming distance (HD) calculations. To quickly identify the minimum value among multiple calculation results, a four-input sense amplifier (SA) is proposed. The performance of the proposed memory is simulated in a 65-nm CMOS process. The post-layout simulation results show good linearity of the multirow read in the SAD and HD modes. The mean time required by the four-input SA to obtain the result is 190 ps. The SAD and HD calculations yield consumptions of 67.44 fJ/byte and 0.64 fJ/bit, respectively, at 0.8 V. Furthermore, a single column-sharing comparator consumes 2.78 and 3.41 pJ at 0.8 V in the SAD and HD modes, respectively.
Yue Zhao 0029, Zhi-Ting Lin, Xiulong Wu, Qiang Zhao 0007, Wenjuan Lu, Chunyu Peng, Zhongzhen Tong, Junning Chen
IEEE Trans. Very Large Scale Integr. Syst.2
2020 Multiple Sharing 7T1R Nonvolatile SRAM With an Improved Read/Write Margin and Reliable Restore Yield
abstract
This article presents a new resistive random access memory (RRAM)-based average 7T1R nonvolatile static random access memory (nvSRAM). This multiple sharing (MS) 7T1R uses MS schemes in which some of the transistors play various roles. Therefore, the MS-7T1R can perform a decoupled read, enhance write capability, and improve the restore yield with a small area cost. Furthermore, the MS-7T1R can offer two alternative modes, i.e., a high speed and a stable mode. Compared with existing technologies, such as the previous 6T-based nvSRAMs, the results show that the proposed architecture provides a remarkable restore yield and ~154% improvement in the read static noise margin (at TT corner and stable mode). In addition, the read delay improves by ~23% (at TT corner and high-speed mode). The write “1” problem of the single bitline is effectively resolved with our proposed write strategy. The static write margin of “1” is improved by ~88.6% compared with the conventional 6T (β = 4) at a power of 1.2 V. In addition, dynamic power is effectively reduced by the use of the single bitline and sub-word-line driver technology.
Zhi-Ting Lin, Yong Wang 0038, Chunyu Peng, Xiulong Wu, Junning Chen
IEEE Trans. Very Large Scale Integr. Syst.1
2020 In-Memory Computing With Double Word Lines and Three Read Ports for Four Operands
abstract
The von Neumann architecture is approaching its limits in terms of scalability and power consumption. In-memory computation is a possible approach to mitigate this limitation. This brief proposes a configurable 8T static random access memory (SRAM) cell with double word lines and three read ports for in-memory computing. In addition to the normal SRAM function, XOR/XNOR and compound Boolean logic operations of three or four operands, such as AND-OR, AND-OR-INVERT, OR-AND, and OR-AND-INVERT, can be performed in one cycle by fully utilizing the three read ports to obtain 13.2-fJ/bit consumption at 0.6 V. The logic operation frequency is 793 MHz at 1.2 V. The proposed SRAM effectively resolves the bottleneck of the existing in-memory computation schemes that only support compound Boolean logic operations with more than two cycles. In addition, the proposed SRAM array scheme can be configured and used as a binary content-addressable memory or a ternary content-addressable memory for searching operations; it achieves 0.24 fJ/search/bit at 0.6 V in the worst case. At 1.2 V, the searching frequency is up to 813 MHz when searching 128 bits with 65-nm technology.
Zhi-Ting Lin, Honglan Zhan, Chunyu Peng, Wenjuan Lu, Xiulong Wu, Junning Chen
IEEE Trans. Very Large Scale Integr. Syst.1
2020 Novel Write-Enhanced and Highly Reliable RHPD-12T SRAM Cells for Space Applications
abstract
In this brief, we proposed, based on the polarity upset mechanism of single-event transient voltage of n-channel metal-oxide-semiconductor (nMOS) transistors, a novel radiation hardened by polar design (RHPD) 12T SRAM cell to enhance the reliability and operation speed for space applications. Simulation results in Semiconductor Manufacturing International Corporation (SMIC) 65-nm CMOS commercial standard process show that the proposed RHPD-12T cell can tolerate all single-node upsets. Meanwhile, compared with We-QUATRO, QUATRO, and dual interlocked storage cell (DICE), the write speed of the proposed cell can be reduced by ~41.8 and ~35.3%, and the static power consumption is reduced by ~41.6 and ~46.3%, respectively. Monte Carlo (MC) simulation has proved that under high frequency and low supply (0.6 V) voltage, RHPD-12T has the minimum write failure probability compared with five other SRAM cells.
Qiang Zhao 0007, Chunyu Peng, Junning Chen, Zhi-Ting Lin, Xiulong Wu
IEEE Trans. Very Large Scale Integr. Syst.4
2019 A review of data sets of short-range wireless networks
Zhi-Ting Lin
Comput. Commun.1
2019 An indexed set representation based multi-objective evolutionary approach for mining diversified top-k high utility patterns
Lei Zhang 0060, Shangshang Yang, Xinpeng Wu, Fan Cheng 0001, Ying Xie 0002, Zhi-Ting Lin
Eng. Appl. Artif. Intell.6
2019 Radiation-Hardened 14T SRAM Bitcell With Speed and Power Optimized for Space Application
abstract
In this paper, a novel radiation-hardened 14-transistor SRAM bitcell with speed and power optimized [radiation-hardened with speed and power optimized (RSP)-14T] for space application is proposed. By circuit- and layout-level optimization design in a 65-nm CMOS technology, the 3-D TCAD mixed-mode simulation results show that the novel structure is provided with increased resilience to single-event upset as well as single-event-multiple-node upsets due to the charge sharing among OFF-transistors. Moreover, the HSPICE simulation results show that the write speed and power consumption of the proposed RSP-14T are improved by ~65% and ~50%, respectively, compared with those of the radiation hardened design (RHD)-12T memory cell.
Chunyu Peng, Jiati Huang, Changyong Liu, Qiang Zhao 0007, Songsong Xiao, Xiulong Wu, Zhi-Ting Lin, Junning Chen, Xuan Zeng 0001
IEEE Trans. Very Large Scale Integr. Syst.7
2018 Clarifying Trust in Social Internet of Things (Extended Abstract)
abstract
Establishing trustworthy relationships among the objects greatly improves the effectiveness of node interaction in the social Internet of Things (IoT). It helps nodes overcome perceptions of uncertainty and risk. However, there are limitations in the existing trust models. In this paper, a comprehensive model of trust is proposed that is tailored to the social IoT. The model includes ingredients such as trustor, trustee, goal, trustworthiness evaluation, decision, action, result, and context. Building on this trust model, we clarify the concepts of trust in the social IoT in five aspects such as: (1) mutuality of trustor and trustee; (2) inferential transfer of trust; (3) transitivity of trust; (4) trustworthiness update; and (5) trustworthiness affected by dynamic environment. With network connectivities that are from real-world social networks, simulations are conducted to evaluate the performance of the social IoT operated with the proposed trust model. An experimental IoT network is used to further validate the proposed trust model.
Zhi-Ting Lin, Liang Dong 0001
ICDE1
2018 Clarifying Trust in Social Internet of Things
abstract
A social approach can be exploited for the Internet of Things (IoT) to manage a large number of connected objects. These objects operate as autonomous agents to request and provide information and services to users. Establishing trustworthy relationships among the objects greatly improves the effectiveness of node interaction in the social IoT and helps nodes overcome perceptions of uncertainty and risk. However, there are limitations in the existing trust models. In this paper, a comprehensive model of trust is proposed that is tailored to the social IoT. The model includes ingredients such as trustor, trustee, goal, trustworthiness evaluation, decision, action, result, and context. Building on this trust model, we clarify the concepts of trust in the social IoT in five aspects such as: 1) mutuality of trustor and trustee; 2) inferential transfer of trust; 3) transitivity of trust; 4) trustworthiness update; and 5) trustworthiness affected by dynamic environment. With network connectivities that are from real-world social networks, a series of simulations are conducted to evaluate the performance of the social IoT operated with the proposed trust model. An experimental IoT network is used to further validate the proposed trust model.
Zhi-Ting Lin, Liang Dong 0001
IEEE Trans. Knowl. Data Eng.1
2018 Average 7T1R Nonvolatile SRAM With R/W Margin Enhanced for Low-Power Application
abstract
A new average 7T1R nonvolatile SRAM for low-power application is presented in this brief, which improves the read and write margin (RM/WM), as well as the restore energy, simply by using the source switch transistor. Simulation results demonstrate that the RM and WM will be improved by ~23% and ~73%, respectively, and the energy consumption will be decreased by ~63% for low-resistance state restoration, compared with the prior art initialization-and-overwrite-7T1R at nMOS typical corner and pMOS typical corner in Taiwan Semiconductor Manufacturing Company's 65-nm technology. In addition, with the column-shared structure, the area penalty is cheerfully acceptable.
Chunyu Peng, Songsong Xiao, Wenjuan Lu, Xiulong Wu, Junning Chen, Zhi-Ting Lin
IEEE Trans. Very Large Scale Integr. Syst.7
2015 Analyzing and modeling mobility for infrastructure-less communication
Zhi-Ting Lin, Xiulong Wu
J. Netw. Comput. Appl.1
2013 Universal scheme improving probabilistic routing in delay-tolerant networks
Zhi-Ting Lin, Xiufang Jiang
Comput. Commun.1
2009 Analysis and design of mobile Wireless Social Model
Zhi-Ting Lin, Yugui Qu, Xiaofang Zhou 0003, Baohua Zhao
Comput. Commun.1
2008 E-Scheme in Delay-Tolerant Networks
Zhi-Ting Lin, Yugui Qu, Baohua Zhao
APNOMS1
2008 An Energy-Efficiency Route Protocol for MIMO-Based Wireless Sensor Networks
Yugui Qu, Zhi-Ting Lin, Baohua Zhao
APNOMS3