VLDB 2026 Research / reviewers in the wild / expert
Kai Ni 0004
dblp:46/3815-4
· DBLP profile ↗
55ranked-venue papers
1as first author
51since 2021 · last 2026
0000-0002-3628-3431ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 55 · 1 first-author · 51 since 2021Software engineering, systems software and programming languages · 8 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | WARP: Workload-Aware Reference Prediction for Reliable Multi-Bit FeFET Readout under Charge-Trapping DegradationabstractFerroelectric FET (FeFET)-based arrays are promising candidates for energy-efficient, high-density non-volatile memory in data-intensive applications. However, charge-trappinginduced degradation and process variations pose significant reliability challenges. These effects lead to reduced memory window and degraded read accuracy over time. We propose a workload-aware degradation modeling and readout framework for FeFET arrays. First, we select a small set of representative workloads to efficiently capture degradation trends across a large workload space. We apply a two-step method to reduce read error: (a) adjust intermediate state currents to widen the separation between states; (b) select optimum reference thresholds based on the shifted distributions. Next, we perform detailed tradeoff analysis involving degradation improvement, on-chip area, and the overhead of a memory-mapped CPU polling system for in-field workload tracking. This is the first work to propose adaptive reference prediction for FeFETs based on runtime workload characteristics. Our framework improves read reliability with minimum hardware overhead and enables scalable in-field monitoring for future FeFET-based systems. Dhruv Thapar, Ashish Reddy Bommana, Arjun Chaudhuri, Kai Ni 0004, Krishnendu Chakrabarty |
ASP-DAC | 4 |
| 2026 | Enhancing Robustness of Content-Addressable Memories for In-Memory Search
Liu Liu 0023, Tomas Sousa Pereira, Mohammad Mehdi Sharifi, Can Li 0024, Kai Ni 0004, Michael T. Niemier, Xiaobo Sharon Hu |
VTS | 7 |
| 2026 | FeFET-Based Analog In-Memory Computing With Inherent Shift-and-Add CapabilityabstractIn-memory computing (IMC) architecture has emerged as a highly promising approach, enhancing the energy efficiency of multiply-and-accumulate (MAC) operations in deep neural networks (DNNs) by embedding parallel computations directly into memory arrays. However, existing ferroelectric FET (FeFET)-based analog IMC designs are often constrained to cell-level optimizations and struggle to achieve high-precision MAC operations. In contrast, high-precision analog IMC architectures typically perform MAC operations for partial inputs and weights within the array in a single cycle and then accumulate partial results over multiple cycles. During this procedure, circuits that handle weight shift-and-add process, whether in digital or analog form, incur significant overhead. This paper presents energy-efficient high-precision analog IMC designs leveraging FeFET technology, which inherently support a shift-and-add mechanism for weights. Initially, we introduce an IMC array paradigm that performs partial MAC operations within each column, and seamlessly incorporates the shift-and-add process for weights by utilizing the analog storage properties of FeFET-based cells. Building upon this paradigm, we propose single-level cell (SLC) FeFET-based designs, namely CurFe and ChgFe, operating in the current and charge modes, respectively. Additionally, to leverage FeFET’s multi-level cell (MLC) properties, we propose a novel hybrid SLC-MLC FeFET-based design, MulFe, which offers higher storage density and energy efficiency. Comprehensive evaluations are conducted at both the circuit and system levels, and the results indicate that the average energy efficiency of the proposed FeFET-based analog IMC designs is 1.32× to 2.71× higher compared to state-of-the-art (SOTA) IMC designs. Qingrong Huang, Yu Qian 0002, Jiahao Cai, Kai Ni 0004, Thomas Kämpfe, Zheyu Yan, Xunzhao Yin, Cheng Zhuo |
IEEE Trans. Computers | 6 |
| 2026 | EvaCAM: A Circuit-Level Evaluation Tool for General Content Addressable MemoriesabstractContent addressable memories (CAMs) are special-purpose in-memory computing units that support parallel searches directly in memory. There is growing interest in CAMs for data-intensive applications such as machine learning, data mining, and bioinformatics, which has led to a rapidly growing CAM design space. CAM cells can be implemented exclusively by CMOS or with various non-volatile memory (NVM) devices. In addition to traditional binary and ternary CAMs (BCAMs and TCAMs), analog CAM (ACAM) and multi-bit CAM (MCAM) designs have recently been introduced, which could further improve density, and also support unique in-memory distance functions. Furthermore, aside from the widely-used exact match function, CAM-based approximate match functions, such as threshold match and best match, have been proposed to further extend the utility of CAMs to new application spaces. As the CAM design space is large, evaluating different CAM design options for a given application is both crucial and challenging. This paper presents EvaCAM, a circuit-level modeling and evaluation tool for CAMs. EvaCAM supports TCAM, ACAM, and MCAM designs implemented in either CMOS or NVMs, for both exact and approximate match functions. It also allows for the exploration of different CAM designs under various optimization targets. EvaCAM has been validated against measured data from fabricated chips and detailed SPICE simulations. A comprehensive design space exploration for CAMs is provided to illustrate the impact of various design decisions and to demonstrate the use cases of EvaCAM. Liu Liu 0023, Mohammad Mehdi Sharifi, Kunshi Wang, Ruibin Mao, Kai Ni 0004, Can Li 0024, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2026 | Modeling and Analysis of Defects and Variations in Multibit FeFET Devices and Crossbar ArchitecturesabstractFerroelectric field-effect transistors (FeFETs) have many promising applications, but the impact of manufacturing imperfections on these devices has yet to be studied comprehensively. We extend a previous FeFET compact model to combine the Preisach ferroelectric capacitor model with the BSIM-SOI MOSFET model. We calibrate this new compact model with data from a technology CAD model that is calibrated against a fabricated metal-ferroelectric-metal capacitor. We analyze polarization defects in the ferroelectric layer using this compact model. We address two classes of defects and map them to stuck-at-fault models, referred to as neutral faults, and stuck-at-plus and stuck-at-minus faults. We also present an analysis of the impact of device-level opens, shorts, coupling faults and extrinsic variations on the multi-level FeFET device in a 1T-1FeFET cell. Simulation results under a clustered fault distribution with a 2% fault rate for faults show a maximum accuracy degradation of 83.86% and 81.66% for ResNet18 and VGG16, respectively when inferencing is performed on faulty FeFET crossbar designs. Similarly, extrinsic variations and clustered polarization defects show a maximum accuracy degradation of 80% and 75% across both DNN models, respectively. We also evaluate the effectiveness of write-verify for mitigating the impact of extrinsic variations, polarization defects, and BEOL faults. Dhruv Thapar, Arjun Chaudhuri, Kai Ni 0004, Krishnendu Chakrabarty |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2025 | PUFiM: A Robust and Efficient FeFET-Based Security Solution Merging Physical Unclonable Function with Compute-in-Memory for Edge AIabstractCompute-in-memory (CiM) has become a promising candidate for edge AI by reducing data movements through insitu operations. However, this emerging computational paradigm also poses the vulnerability of model leakage as the weights are stored in plaintext for computing. While prior works have explored lightweight encryption methods, CiM is usually considered a separate module instead of a system component, leaving the origin of keys unclear and unprotected. Physical unclonable functions (PUFs) offer a potential origin of keys, but a comprehensive framework for securing key generation and delivery remains lacking. Besides, the complementary ciphertext storage incurs substantial costs and degrades the performance. This work proposes PUFiM, a robust and efficient security solution for edge computing based on ferroelectric FETs (FeFETs). For the first time, a strong PUF is synergized with CiM to enable authentication, key generation, and encrypted computations within a unified array for comprehensive protection. To achieve this synergization, a high-density hybrid storage and computation approach combining PUF and weight bits via multi-level cell (MLC) FeFETs is proposed. Besides, two PUF enhancement techniques and a novel mapping scheme are developed to improve security and efficiency further. Results show that PUFiM withstands PUF modeling attacks with up to $\mathbf{1 0 M}$ samples. Moreover, PUFiM reduces the inference accuracy by $\gt 60 \%$ under 95% key leakage and achieves $\gt 9.7 \times$ compute density and $\gt 1.2 \times$ energy efficiency improvement compared with the state-of-the-art SRAM/NVM secure CiMs. Taixin Li, Thomas Kämpfe, Kai Ni 0004, Narayanan Vijaykrishnan, Huazhong Yang, Xueqing Li 0002 |
DAC | 4 |
| 2025 | Towards Uncertainty-aware Robotic Perception via Mixed-signal BNN Engine Leveraging Probabilistic Quantum TunnelingabstractIntegrating deep learning with environmental perception enhances robotic adaptability to complex tasks. However, its “black-box” nature, such as the lack of uncertainty quantification, poses challenges for safety-critical applications, particularly in unstructured and noisy environments. Bayesian neural networks (BNNs) offer uncertainty quantification but are limited by high hardware overhead, restricting real-time implementation on resource-constrained robots. This paper presents a mixedsignal hardware accelerator for BNNs, utilizing probabilistic quantum tunneling in fully depleted silicon-on-insulator (FDSOI) transistors to enable efficient, real-time uncertainty quantification. Device measurements indicate high-quality Gaussian random variable generation, validated through quantile-quantile plot analysis, with a high correlation coefficient ($r=0.997$) at $200 \mathrm{fJ} /$ sample. Leveraging such compact randomness, the parallel architecture achieved $10^{3}-10^{4} \times$ latency reduction at less than $2 \times$ area cost. Finally, in uncertainty-aware visual localization application of autonomous underwater vehicles, the BNN model effectively distinguishes data noise from model uncertainty, yielding significant information gain and enhancing the resampling efficiency by $4.5 \times$ at same accuracy. Likai Pei, Xingtian Wang, Xueji Zhao, Wanxin Huang, Boyang Cheng, Halid Mulaosmanovic, Stefan Dünkel, Dominik Kleimaier, Sven Beyer, Kai Ni 0004, Mengxue Hou, Michael T. Niemier, Ningyuan Cao |
DAC | 11 |
| 2025 | NVCiM-PT: An NVCiM-Assisted Prompt Tuning Framework for Edge LLMsabstractLarge Language Models (LLMs) deployed on edge devices, known as edge LLMs, need to continuously fine-tune their model parameters from user-generated data under limited resource constraints. However, most existing learning methods are not applicable for edge LLMs because of their reliance on high resources and low learning capacity. Prompt tuning (PT) has recently emerged as an effective fine-tuning method for edge LLMs by only modifying a small portion of LLM parameters, but it suffers from user domain shifts, resulting in repetitive training and losing resource efficiency. Conventional techniques to address domain shift issues often involve complex neural networks and sophisticated training, which are incompatible for PT for edge LLMs. Therefore, an open research question is how to address domain shift issues for edge LLMs with limited resources. In this paper, we propose a prompt tuning framework for edge LLMs, exploiting the benefits offered by non-volatile computing-in-memory (NVCiM) architectures. We introduce a novel NVCiM-assisted PT framework, where we narrow down the core operations to matrix-matrix multiplication, which can then be accelerated by performing in-situ computation on NVCiM. To the best of our knowledge, this is the first work employing NVCiM to improve the edge LLM PT performance. Ruiyang Qin, Zheyu Yan, Liu Liu 0023, Dancheng Liu, Amir Nassereldine, Jinjun Xiong, Kai Ni 0004, Xiaobo Sharon Hu, Yiyu Shi 0001 |
DATE | 8 |
| 2025 | STAMP-2.5D: Structural and Thermal Aware Methodology for Placement in 2.5D IntegrationabstractChiplet-based architectures and advanced packaging have emerged as transformative approaches in semiconductor design. While conventional physical design for 2.5D heterogeneous systems typically prioritizes wirelength reduction through tight chiplet packing, this strategy introduces thermal bottlenecks and intensifies coefficient of thermal expansion (CTE) mismatches, compromising long-term reliability. Addressing these challenges requires holistic consideration of thermal performance, mechanical stress, and interconnect efficiency. We introduce STAMP-2.5D: Structural and Thermal Aware Methodology for Placement in 2.5D integration, the first automated placement methodology that simultaneously optimizes these critical factors. Our approach employs finite element analysis to simulate temperature distributions and stress profiles across chiplet configurations while minimizing interconnect wirelength. Experimental results demonstrate that a thermal and structurally aware automated placement approach reduces overall stress by 11 %, maintains excellent thermal performance with a negligible$\mathbf{0. 5 \%}$temperature increase and simultaneously reduces total wirelength by 11 % compared to temperature-only optimization. Additionally, we conduct an exploratory study on the effects of temperature gradients on structural integrity, providing crucial insights for reliability-conscious chiplet design. Varun Darshana Parekh, Zachary Wyatt Hazenstab, Srivatsa Rangachar Srinivasa, Krishnendu Chakrabarty, Kai Ni 0004, Narayanan Vijaykrishnan |
ICCD | 5 |
| 2025 | FeTest: Defect Analysis and March Test Solution for FeFETs *abstractFerroelectric Field-Effect Transistors (FeFETs) are emerging as promising candidates for non-volatile memory and in-memory computing due to their low power consumption, fast switching speeds, and high integration density. However, device-level defects and process variations pose significant challenges to their reliability, particularly for multi-level cell (MLC) FeFETs. Polarization defects in the ferroelectric layer and back-end-of-line (BEOL) defects lead to memory window reduction and erroneous computations in FeFET-based crossbars. We propose a March test solution tailored for MLC FeFETs that detects and diagnoses BEOL defects in the presence of underlying process variations and polarization defects. Dhruv Thapar, Arjun Chaudhuri, Kai Ni 0004, Krishnendu Chakrabarty |
ITC | 3 |
| 2025 | From Secure Storage to Compute-in-Memory: A Versatile Memory System using 1T-nC FeRAMabstractThis study explores the versatility of 1T-nC ferro-electric RAM (FeRAM) memory capable of operating in both secure storage and highly energy-efficient Compute-in-Memory (CiM) modes without incurring any extra design overhead. In contrast to conventional, volatile 1T-1C Dynamic Random Access Memory (DRAM), having significant energy cost due to the continuous refresh cycles, FeRAM presents a robust non-volatile memory solution, enabling efficient data retention with minimal power consumption. Bitwise logic operations (AND, OR) are efficiently implemented using a single, unmodified 1T-nC FeRAM cell and single-row activation (SRA), while NOT operation is realized by a modified sense amplifier. These facilitate bulk bitwise operations with significant energy benefits compared to multi-row based DRAM solutions. Furthermore, we demonstrate key-based (K) cryptographic operations which help in storing Plain Text (PT) in Cipher Text (CT) format tightly coupled with the memory write/read cycles in secure-mode operation with minimal overload while maintaining robust security. In addition, the advantages of vertical 3D integration of $1 \mathrm{~T}-\mathrm{nC}$ FeRAM are evaluated, and thermal profiling of the remnant polarization ($\mathrm{P}_{\mathrm{R}}$) is done. Furthermore, we demonstrate that 1T-nC FeRAM memory reduces energy consumption by 31% compared to DRAM by assessing 8 real-world data-intensive workloads utilizing bulk bitwise operations. Rakesh Acharya, Rudra Biswas, Jiahui Duan, Prapti Panigrahi, Kai Ni 0004, Narayanan Vijaykrishnan |
VLSI-SoC | 5 |
| 2025 | High-Performance In-Memory Bayesian Inference With Multi-Bit Ferroelectric FETabstractConventional neural network-based machine learning algorithms often encounter difficulties in data-limited scenarios or where interpretability is critical. Conversely, Bayesian inference-based models excel with reliable uncertainty estimates and explainable predictions. Recently, many in-memory computing (IMC) architectures achieve exceptional computing capacity and efficiency for neural network tasks leveraging emerging nonvolatile memory (NVM) technologies. However, their application in Bayesian inference remains limited because the operations in Bayesian inference differ substantially from those in neural networks. In this article, we introduce a compact in-memory Bayesian inference engine with high efficiency and performance utilizing a multi-bit ferroelectric field-effect transistor (FeFET). This design encodes a Bayesian model within a compact FeFETbased crossbar by mapping quantized probabilities to discrete FeFET states. Consequently, the crossbar’s outputs naturally represent the output posteriors of the Bayesian model. Our design facilitates efficient Bayesian inference, accommodating various input types and probability precisions, without additional calculation circuitry. As the first FeFET-based in-memory Bayesian inference engine, our design demonstrates a notable storage density of 26.32 Mb/mm2and a computing efficiency of 581.40 TOPS/W in a representative Bayesian classification task, indicating a 10.7×/43.4× compactness/efficiency improvement compared to the state-of-the-art alternative. Utilizing the proposed Bayesian inference engine, we develop a feature selection system that efficiently addresses a representative NP-hard optimization problem, showcasing our design’s capability and potential to enhance various Bayesian inference-based applications. Test results suggest that our design identifies the essential features, enhancing the model’s performance while reducing its complexity, surpassing the latest implementation in operation speed and algorithm efficiency by 2.9×/2.0×, respectively. Chao Li 0065, Xuchu Huang, Ruibin Mao, Thomas Kämpfe, Kai Ni 0004, Can Li 0024, Xunzhao Yin, Cheng Zhuo |
IEEE Trans. Computers | 8 |
| 2025 | A Scalable 2T-1FeFET-Based Content Addressable Memory Design for Energy Efficient Data SearchabstractContent addressable memory (CAM) is widely used in advanced machine learning models and data-intensive applications for associative search tasks, thanks to the highly parallel pattern matching capability. Most state-of-the-art CAM designs primarily aim to reduce the CAM cell area by utilizing nonvolatile memories (NVMs). However, there has been limited research on optimizing the design and energy efficiency of NVM-based CAMs for practical deployment in edge devices and AI hardware. This article introduces a general compact and energy efficient CAM design scheme that minimizes design overhead by using only one NVM device per cell. Our proposed CAM design realizes both binary CAM (BCAM) and multibit CAM (MCAM) by leveraging the binary and multilevel storage property of NVM devices without additional cell overheads. Additionally, we propose an adaptive matchline (ML) precharge and discharge scheme to further optimize search energy by significantly reducing the ML voltage swing. Ferroelectric field-effect transistors (FeFETs) serve as representative NVMs in our proposed design, and we present a 2T-1FeFET CAM array incorporating a sense amplifier that implements the proposed ML scheme. Evaluation results show that our proposed 2T-1FeFET BCAM design achieves energy efficiency improvements of$6.64\times $/$4.74\times $/$9.14\times $/$3.02\times $compared to CMOS/ReRAM/STT-MRAM/2FeFET BCAM arrays, while 2T-1FeFET MCAM design achieves$8.25\times $/$5.68\times $/$56.35\times $better-energy efficiency compared to ReRAM/3T-1FeFET/1FeFET-1R MACM arrays. Benchmarking results demonstrate that our BCAM/MCAM approach provides$3.2\times $/$3.7\times $and$2.0\times $/$2.2\times $energy-delay product improvement over the 2T-2R and 2FeFET CAM in accelerating query processing applications. Jiahao Cai, Hamza Errahmouni Barkam, Mohsen Imani, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | CSA-CiM: Enhancing Multifunctional Computing-in-Memory With Configurable Sense AmplifiersabstractComputing-in-memory (CiM) effectively alleviates the memory wall problem faced by traditional von Neumann architectures when handling data-intensive applications. Most CiM arrays employ dedicated sense amplifiers (SAs) to perform specific functions, and prior configurable CiM arrays achieve multifunctionality by stacking multiple SAs with corresponding functions. However, the independent nature of these SAs, particularly the analog-to-digital converter (ADC), results in excessive energy and area consumption. In this article, we propose a configurable multifunctional ferroelectric field effect transistor (FeFET)-based CiM array design, including configurable peripheral circuit with corresponding multifunctionalities and reusable SA components, to reduce energy consumption and latency. The array cells perform logical AND and XNOR operations, and the proposed SA can be configured to operate in either ADC or winner-take-all (WTA) modes, thereby enabling the array to implement both multiplication-accumulation (MAC) and associative search operations. Instead of operating independently, the WTA component within the SA participates as a flash stage in successive approximation register (SAR) conversions in ADC mode, thus enhancing the WTA utilization, energy efficiency and compactness. By integrating the multifunctional CiM array and the configurable SA, our design supports MAC, Hamming-distance computation (HDC), and nearest neighbor search (NNS) operations within the same structure. Compared to existing works, our design achieves energy efficiency improvements of$7.2\times $for MAC,$2.9\times $for HDC, and EDP improvement of$6.4\times $for NNS, respectively. Yuxiao Jiang, Kai Ni 0004, Thomas Kämpfe, Cheng Zhuo, Zheyu Yan, Xunzhao Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | A Homogeneous FeFET-Based Time-Domain Compute-in-Memory Fabric for Matrix-Vector Multiplication and Associative SearchabstractMatrix-vector multiplication (MVM) and content-based search are two key operations in many machine learning workloads. This article proposes a ferroelectric FET (FeFET) time-domain compute-in-memory (TD-CiM) array that can accelerate both operations in a homogeneous fabric. We demonstrate that 1) the AND and xor/XNOR logic functions required by MVM and content-based search can be realized using a single compute-in-memory (CiM) cell composed of 2FeFETs connected in series; 2) an inverter chain-based TD-CiM array along with a two-phase time-domain computation principle of the TD-CiM can be employed to implement the MVM and content-based search functions; 3) a signal delay-to-digital output conversion can be implemented by associating a loading capacitor with each stage of the inverter chain-based TD-CiM array, ensuring the full digital compatibility; and 4) the proposed 2FeFET cell and inverter chain-based TD-CiM array are robust against FeFET variation according to our comprehensive theoretical and experimental validation. We show how the FeFET TD-CiM can be exploited to accelerate hyperdimensional computing (HDC) and adjusted to process different tasks through dynamic and fine-grained resource allocation. HDC application benchmarking results show that the proposed FeFET-based TD-CiM offers on average$106\times $/$63\times $energy reduction/speedup compared to GPU-based implementation. With more than 8500 TOPS/W energy-efficiency, the proposed FeFET-based TD-CiM exhibits huge potential as a processing fabric for various memory-intensive applications. Xunzhao Yin, Qingrong Huang, Hamza Errahmouni Barkam, Franz Müller 0001, Shan Deng, Alptekin Vardar, Sourav De 0002, Zhouhang Jiang, Mohsen Imani, Ulf Schlichtmann, Xiaobo Sharon Hu, Cheng Zhuo, Thomas Kämpfe, Kai Ni 0004 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 14 |
| 2025 | Ferroelectric Compute-in-Memory Framework for Solving Pure and Mixed Strategy Nash EquilibriumabstractNash equilibrium (NE) is a key concept in game theory, but verifying its existence is NP-complete. Recent advancements proposed quantum NE solvers that identify pure strategy NE solutions (binary solutions) by integrating slack terms into the objective function, known as slack-quadratic unconstrained binary optimization (S-QUBO). However, S-QUBO alters the objective function and can lead to incorrect solutions. Additionally, current solvers only find a limited number of pure strategy NE solutions and cannot address mixed strategy NE (decimal solutions), leaving many solutions unexplored. In this work, we propose C-Nash, a novel ferroelectric compute-in-memory (CiM) framework capable of efficiently addressing both pure and mixed strategy NE solutions. C-Nash consists of 1) a transformation method that transforms quadratic optimization into a MAX-QUBO form without incorporating additional slack variables, thus avoiding objective function changes; 2) A ferroelectric FET (FeFET) based CiM bi-crossbar structure and winner-takes-all (WTA) tree for accelerating the MAX-QUBO form in a single iteration; 3) An efficient operation flow including a rank-based QUBO reformulation algorithm that simplifies the QUBO matrices to reduce hardware overhead, and a two-phase based simulated annealing (SA) logic for finding NE solutions; 4) A FeFET-based crossbar macro for experimental demonstration. Experimental results show that C-Nash increases the success rate for identifying NE solutions by 68.6% while saving$3\times $in chip size. Furthermore, C-Nash can find all pure and mixed NE solutions, unlike D-Wave based quantum approaches which only find some pure strategy NE solutions. Additionally, C-Nash significantly reduces the time-to-solution by up to$157.9\times $/$79.0\times $compared to D-Wave 2000 Q6 and D-Wave Advantage 4.1, respectively. Yu Qian 0002, Ding Huang, Alptekin Vardar, Nellie Laleni, Kai Ni 0004, Thomas Kämpfe, Cheng Zhuo, Xunzhao Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2024 | FeBiM: Efficient and Compact Bayesian Inference Engine Empowered with Ferroelectric In-Memory ComputingabstractIn scenarios with limited training data or where explainability is crucial, conventional neural network-based machine learning models often face challenges. In contrast, Bayesian inference-based algorithms excel in providing interpretable predictions and reliable uncertainty estimation in these scenarios. While many state-of-the-art in-memory computing (IMC) architectures leverage emerging non-volatile memory (NVM) technologies to offer unparalleled computing capacity and energy efficiency for neural network workloads, their application in Bayesian inference is limited. This is because the core operations in Bayesian inference, i.e., cumulative multiplications of prior and likelihood probabilities, differ significantly from the multiplication-accumulation (MAC) operations common in neural networks, rendering them generally unsuitable for direct implementation in most existing IMC designs. In this paper, we propose FeBiM, an efficient and compact Bayesian inference engine powered by multi-bit ferroelectric field-effect transistor (FeFET)-based IMC. FeBiM effectively encodes the trained probabilities of a Bayesian inference model within a compact FeFET-based crossbar. It maps quantized logarithmic probabilities to discrete FeFET states. As a result, the accumulated outputs of the crossbar naturally represent the posterior probabilities, i.e., the Bayesian inference model's output given a set of observations. This approach enables efficient in-memory Bayesian inference without the need for additional calculation circuitry. As the first FeFET-based in-memory Bayesian inference engine, FeBiM achieves an impressive storage density of 26.32 Mb/mm2 and a computing efficiency of 581.40 TOPS/W in a representative Bayesian classification task. These results demonstrate 10.7×/43.4× improvement in compactness/efficiency compared to the state-of-the-art hardware implementation of Bayesian inference. Chao Li 0065, Ruibin Mao, Can Li 0024, Thomas Kämpfe, Kai Ni 0004, Xunzhao Yin |
DAC | 7 |
| 2024 | C-Nash: A Novel Ferroelectric Computing-in-Memory Architecture for Solving Mixed Strategy Nash EquilibriumabstractThe concept of Nash equilibrium (NE), pivotal within game theory, has garnered widespread attention across numerous industries. However, verifying the existence of NE poses a significant computational challenge, classified as an NP-complete problem. Recent advancements introduced several quantum Nash solvers aimed at identifying pure strategy NE solutions (i.e., binary solutions) by integrating slack terms into the objective function, commonly referred to as slack-quadratic unconstrained binary optimization (S-QUBO). However, incorporation of slack terms into the quadratic optimization results in changes of the objective function, which may cause incorrect solutions. Furthermore, these quantum solvers only identify a limited subset of pure strategy NE solutions, and fail to address mixed strategy NE (i.e., decimal solutions), leaving many solutions undiscovered. In this work, we propose C-Nash, a novel ferroelectric computing-in-memory (CiM) architecture that can efficiently handle both pure and mixed strategy NE solutions. The proposed architecture consists of (i) a transformation method that converts quadratic optimization into a MAX-QUBO form without introducing additional slack variables, thereby avoiding objective function changes; (ii) a ferroelectric FET (FeFET) based bi-crossbar structure for storing payoff matrices and accelerating the core vector-matrix-vector (VMV) multiplications of QUBO form; (iii) A winner-takes-all (WTA) tree implementing the MAX form and a two-phase based simulated annealing (SA) logic for searching NE solutions. Evaluations show that C-Nash has up to 68.6% increase in the success rate for identifying NE solutions, finding all pure and mixed NE solutions rather than only a portion of pure NE solutions, compared to D-Wave based quantum approaches. Moreover, C-Nash boasts a reduction up to 157.9X/79.0X in time-to-solutions compared to D-Wave 2000 Q6 and D-Wave Advantage 4.1, respectively. Yu Qian 0002, Kai Ni 0004, Thomas Kämpfe, Cheng Zhuo, Xunzhao Yin |
DAC | 2 |
| 2024 | HyCiM: A Hybrid Computing-in-Memory QUBO Solver for General Combinatorial Optimization Problems with Inequality ConstraintsabstractComputationally challenging combinatorial optimization problems (COPs) play a fundamental role in various applications. To tackle COPs, many Ising machines and Quadratic Unconstrained Binary Optimization (QUBO) solvers have been proposed, which typically involve direct transformation of COPs into Ising models or equivalent QUBO forms (D-QUBO). However, when addressing COPs with inequality constraints, this D-QUBO approach introduces numerous extra auxiliary variables, resulting in a substantially larger search space, increased hardware costs, and reduced solving efficiency. In this work, we propose HyCiM, a novel hybrid computing-inmemory (CiM) based QUBO solver framework, designed to overcome aforementioned challenges. The proposed framework consists of (i) an innovative transformation method (first to our known) that converts COPs with inequality constraints into an inequality-QUBO form, thus eliminating the need of expensive auxiliary variables and associated calculations; (ii) "inequality filter", a ferroelectric FET (FeFET)-based CiM circuit that accelerates the inequality evaluation, and filters out infeasible input configurations; (iii) a FeFET-based CiM annealer that is capable of approaching global solutions of COPs via iterative QUBO computations within a simulated annealing process. The evaluation results show that HyCiM drastically narrows down the search space, eliminating 2100 to 22536 infeasible input configurations compared to the conventional D-QUBO approach. Consequently, the narrowed search space, reduced to 2100 feasible input configurations, leads to a substantial hardware area overhead reduction, ranging from 88.06% to 99.96%. Additionally, HyCiM consistently exhibits a high solving efficiency, achieving a remarkable average success rate of 98.54%, whereas D-QUBO implementatoin shows only 10.75%. Yu Qian 0002, Kai Ni 0004, Alptekin Vardar, Thomas Kämpfe, Xunzhao Yin |
DAC | 3 |
| 2024 | Energy Efficient Dual Designs of FeFET-Based Analog In-Memory Computing with Inherent Shift-Add CapabilityabstractIn-memory computing (IMC) architecture emerges as a promising paradigm, improving the energy efficiency of multiply-and-accumulate (MAC) operations within deep neural networks (DNNs) by integrating the parallel computations within the memory arrays. Various high-precision analog IMC array designs have been developed based on both SRAM and emerging non-volatile memories (NVMs). These designs perform MAC operations of partial input and weight, with the corresponding partial products then fed into shift-add circuitry to produce the final MAC results. However, existing works often present intricate shift-add process for weight. The traditional digital shift-add process is limited in throughput due to time-multiplexing of ADCs, and advancing the shift-add process to the analog domain necessitates customized circuit implementations, resulting in compromises in energy and area efficiency. Furthermore, the joint optimization of the partial MAC operations and the weight shift-add process is rarely explored. In this paper, we propose novel, energy efficient dual designs of ferroelectric FET (FeFET) based high precision analog IMC featuring inherent shift-add capability. We introduce a FeFET based IMC paradigm that performs partial MAC in each column, and inherently integrates the shift-add process for 4-bit weights by leveraging FeFET's analog storage characteristics. This paradigm supports both 2's complement mode (2CM) and non-2's complement mode (N2CM) MAC, thereby offering flexible support for 4-/8-bit weight data in 2's complement format. Building upon this paradigm, we propose novel FeFET based dual designs, CurFe for the current mode and ChgFe for the charge mode, to accommodate the high precision analog domain IMC architecture. Evaluation results at circuit and system levels indicate that the circuit/system-level energy efficiency of the proposed FeFET-based analog IMC is 1.56×/1.37× higher when compared to the state-of-the-art analog IMC designs. Qingrong Huang, Yu Qian 0002, Kai Ni 0004, Thomas Kämpfe, Xunzhao Yin |
DAC | 4 |
| 2024 | A FeFET-based Time-Domain Associative Memory for Multi-bit Similarity ComputationabstractThe exponential growth of data across various domains of human society necessitates the rapid and efficient data processing. In many contemporary data-intensive applications, similarity computation (SC) is one of the most fundamental and indispensable operations. In recent years, In-memory computing (IMC) architectures have been designed to accelerate SC by reducing data movement costs, however, they encounter challenges with signal domain conversion, variation sensitivity, and limited precision. This paper proposes a ferroelectric FET (FeFET) based time-domain (TD) associative memory (AM) for energy efficient SC. Such TD design can convert its output (i.e., time interval) to digits with relatively simple sensing circuitry thus saves large amount of area and energy compared with conventional IMC designs that process analog voltage/current signals. The variable-capacitance (VC) delay chain structure in our design supports quantitative SC and enhances robustness against variations. Furthermore, by exploiting multi-domain ferroelctric FET (FeFET), our design is capable of performing SC on vectors with multi-bit element, enabling support for higher-precision algorithms. Simulation results show that the proposed TD-AM achieves 13.8x/1.47x energy saving of our design compared to CMOS/NVM based TD-IMC designs. Additionally, our design exhibits good robustness in monte carlo simulation with variation extracted from experimental measurements. Investigation on precision of hyperdimensional computing (HDC) show that higher element precision reduces the size of HDC model when considering to achieve same accuracy, indicating an improved efficiency. Benchmarkings against GPU demonstrate in general 2/3 orders of magnitude speedup/energy efficiency improvement of our design. Our proposed multi-bit TD-AM promises energy-efficient quantitative SC for diverse intensive data processing application, especially in energy-constrained scenarios. Qingrong Huang, Hamza Errahmouni Barkam, Jianyi Yang 0003, Thomas Kämpfe, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Mohsen Imani, Cheng Zhuo, Xunzhao Yin |
DATE | 6 |
| 2024 | CafeHD: A Charge-Domain FeFET-Based Compute-in-Memory Hyperdimensional Encoder with Hypervector MergingabstractHyperdimensional computing (HDC) is an emerging paradigm that employs hypervectors (HV s) to emulate cognitive tasks. In HDC, the most time-consuming and power-hungry process is encoding, the first step that maps raw data into HV s. There have been non-volatile memory (NVM) based computing-in-memory (CiM) HDC encoding designs, which exploit the intrinsic HDC characteristics of high parallelism, massive data, and robustness. These NVM-based CiMs have shown great potential in reducing encoding time and power consumption. Among them, the ferroelectric field-effect transistor (FeFET) based designs show ultra-high energy efficiency. However, existing FeFET-based HDC encoding designs face the challenges of energy -consuming current-mode addition, inefficient HV storage, limited endurance, and single encoding method support. These challenges limit the energy efficiency, lifetime, and versatility of the designs. This work proposes an energy-efficient charge-domain FeFET-based in-memory HDC encoder, i.e., CafeHD, with extended lifetime, good versatility, and comparable accuracy. Area-efficient charge-domain computing is proposed in HDC encoding for the first time, which enables CafeHD with ultra-low power and high scalability. An HV merging technique is explored to improve the performance. A low-cost partial MAJ interface is also proposed to reduce writes. Besides, CafeHD also supports two widely used encoding methods. Results show that CafeHD on average achieves 10.9×/12.7×/3.5× speedup and 103.3×/21.9×/6.3× energy effi-ciency with ~84 % write times reduction and similar accuracy compared with the state-of-the-art ReRAM/PCMlFeFET-based CiM design for HDC encoding, respectively. Taixin Li, Hongtao Zhong, Juejian Wu, Thomas Kämpfe, Kai Ni 0004, Narayanan Vijaykrishnan, Huazhong Yang, Xueqing Li 0002 |
DATE | 5 |
| 2024 | Reconfigurable Frequency Multipliers Based on Complementary Ferroelectric TransistorsabstractFrequency multipliers, a class of essential electronic components, play a pivotal role in contemporary signal processing and communication systems. They serve as crucial building blocks for generating high-frequency signals by multiplying the frequency of an input signal. However, traditional frequency multipliers that rely on nonlinear devices often require energy- and area-consuming filtering and amplification circuits, and emerging designs based on an ambipolar ferroelectric transistor require costly non-trivial characteristic tuning or complex technology process. In this paper, we show that a pair of standard ferroelectric field effect transistors (FeFETs) can be used to build compact frequency multipliers without aforementioned technology issues. By leveraging the tunable parabolic shape of the 2FeFET structures' transfer characteristics, we propose four reconfigurable frequency multipliers, which can switch between signal transmission and frequency doubling. Furthermore, based on the 2FeFET structures, we propose four frequency multipliers that realize triple, quadruple frequency modes, elucidating a scalable methodology to generate more multiplication harmonics of the input frequency. Performance metrics such as maximum operating frequency, power, etc., are evaluated and compared with existing works. We also implement a practical case of frequency modulation scheme based on the proposed reconfigurable multipliers without additional devices. Our work provides a novel path of scalable and reconfigurable frequency multiplier designs based on devices that have characteristics similar to FeFETs, and show that FeFETs are a promising candidate for signal processing and communication systems in terms of maximum operating frequency and power. Jianyi Yang 0003, Cheng Zhuo, Thomas Kämpfe, Kai Ni 0004, Xunzhao Yin |
DATE | 5 |
| 2024 | Low Power and Temperature- Resilient Compute-In-Memory Based on Subthreshold-FeFETabstractCompute-in-memory (CiM) is a promising solution for addressing the challenges of artificial intelligence (AI) and the Internet of Things (IoT) hardware such as “memory wall” issue. Specifically, CiM employing nonvolatile memory (NVM) devices in a crossbar structure can efficiently accelerate multiply-accumulation (MAC) computation, a crucial operator in neural networks among various AI models. Low power CiM designs are thus highly desired for further energy efficiency optimization on AI models. Ferroelectric FET (FeFET), an emerging device, is attractive for building ultra-low power CiM array due to CMOS compatibility, high ION /$I$O F F ratio, etc. Recent studies have explored FeFET based CiM designs that achieve low power consumption. Nevertheless, subthreshold-operated FeFETs, where the operating voltages are scaled down to the subthreshold region to reduce array power consumption, are particularly vulnerable to temperature drift, leading to accuracy degradation. To address this challenge, we propose a temperature-resilient 2T-1FeFET CiM design that performs MAC operations reliably at subthreahold region from 0°C to 85°C, while consuming ultra-low power. Benchmarked against the VGG neural network architecture running the CIFAR-10 dataset, the proposed 2T1FeFET CiM design achieves 89.45% CIFAR-10 test accuracy. Compared to previous FeFET based CiM designs, it exhibits immunity to temperature drift at an 8-bit wordlength scale, and achieves better energy efficiency with 2866 TOPS/W. Xuchu Huang, Jianyi Yang 0003, Kai Ni 0004, Hussam Amrouch, Cheng Zhuo, Xunzhao Yin |
DATE | 4 |
| 2024 | REMNA: Variation-Resilient and Energy-Efficient MLC FeFET Computing-in-Memory Using NAND Flash-Like Read and Adaptive ControlabstractNonvolatile memory (NVM)-based computing-in-memory (CiM) has shown promising prospects in deep neural network (DNN) inference at the edge thanks to its nonvolatility and high density. Moreover, most NVMs support multi-level cell (MLC) storage, which can further boost energy efficiency and storage density. However, MLC NVM-based CiMs suffer from degraded accuracy due to device nonidealities, including large variations, nonlinear current distribution, and state drifts. Although prior works have explored various mitigation measures, such as hybrid SLC/MLC, write-and-verify, and local recovery units, the substantial costs from software support, energy, latency, and area still limit the performance. Therefore, the tradeoff between inference accuracy, storage density and compute density has become a vital challenge in NVM-based CiMs. Taixin Li, Hongtao Zhong, Yixin Xu 0001, Narayanan Vijaykrishnan, Kai Ni 0004, Huazhong Yang, Thomas Kämpfe, Xueqing Li 0002 |
ICCAD | 5 |
| 2024 | TAP-CAM: A Tunable Approximate Matching Engine based on Ferroelectric Content Addressable MemoryabstractPattern search is crucial in numerous analytic applications for retrieving data entries akin to the query. Content Addressable Memories (CAMs), an in-memory computing fabric, directly compare input queries with stored entries through embedded comparison logic, facilitating fast parallel pattern search in memory. While conventional CAM designs offer exact match functionality, they are inadequate for meeting the approximate search needs of emerging data-intensive applications. Some recent CAM designs propose approximate matching functions, but they face limitations such as excessively large cell area or the inability to precisely control the degree of approximation. In this paper, we propose TAP-CAM, a novel ferroelectric field effect transistor (FeFET) based ternary CAM (TCAM) capable of both exact and tunable approximate matching. TAP-CAM employs a compact 2FeFET-2R cell structure as the entry storage unit, and similarities in Hamming distances between input queries and stored entries are measured using an evaluation transistor associated with the matchline of CAM array. The operation, robustness and performance of the proposed design at array level have been discussed and evaluated, respectively. We conduct a case study of K-nearest neighbor (KNN) search to benchmark the proposed TAP-CAM at application level. Results demonstrate that compared to 16T CMOS CAM with exact match functionality, TAP-CAM achieves a 16.95× energy improvement, along with a 3.06% accuracy enhancement. Compared to 2FeFET TCAM with approximate match functionality, TAP-CAM achieves a 6.78× energy improvement. Chenyu Ni, Che-Kai Liu, Liu Liu 0023, Mohsen Imani, Thomas Kämpfe, Kai Ni 0004, Michael T. Niemier, Xiaobo Sharon Hu, Cheng Zhuo, Xunzhao Yin |
ICCAD | 7 |
| 2024 | Robust Implementation of Retrieval-Augmented Generation on Edge-based Computing-in-Memory ArchitecturesabstractLarge Language Models (LLMs) deployed on edge devices learn through fine-tuning and updating a certain portion of their parameters. Although such learning methods can be optimized to reduce resource utilization, the overall required resources remain a heavy burden on edge devices. Instead, Retrieval-Augmented Generation (RAG), a resource-efficient LLM learning method, can improve the quality of the LLM-generated content without updating model parameters. However, the RAG-based LLM may involve repetitive searches on the profile data in every user-LLM interaction. This search can lead to significant latency along with the accumulation of user data. Conventional efforts to decrease latency result in restricting the size of saved user data, thus reducing the scalability of RAG as user data continuously grows. It remains an open question: how to free RAG from the constraints of latency and scalability on edge devices? In this paper, we propose a novel framework to accelerate RAG via Computing-in-Memory (CiM) architectures. It accelerates matrix multiplications by performing in-situ computation inside the memory while avoiding the expensive data transfer between the computing unit and memory. Our framework, Robust CiM-backed RAG (RoCR), utilizing a novel contrastive learning-based training method and noise-aware training, can enable RAG to efficiently search profile data with CiM. To the best of our knowledge, this is the first work utilizing CiM to accelerate RAG. Ruiyang Qin, Zheyu Yan, Dewen Zeng, Zhenge Jia, Dancheng Liu, Ahmed Abbasi, Zhi Zheng 0002, Ningyuan Cao, Kai Ni 0004, Jinjun Xiong, Yiyu Shi 0001 |
ICCAD | 10 |
| 2024 | TReCiM: Lower Power and Temperature-Resilient Multibit 2FeFET-1T Compute-in-Memory DesignabstractCompute-in-memory (CiM) emerges as a promising solution to solve hardware challenges in artificial intelligence (AI) and the Internet of Things (IoT), particularly addressing the "memory wall" issue. By utilizing nonvolatile memory (NVM) devices in a crossbar structure, CiM efficiently accelerates multiplyaccumulate (MAC) computations, the crucial operations in neural networks and other AI models. Among various NVM devices, Ferroelectric FET (FeFET) is particularly appealing for ultra-low-power CiM arrays due to its CMOS compatibility, voltage-driven write/read mechanisms and high ION/IOFF ratio. Moreover, subthreshold-operated FeFETs, which operate at scaling voltages in the subthreshold region, can further minimize the power consumption of CiM array. However, subthreshold-FeFETs are susceptible to temperature drift, resulting in computation accuracy degradation. Existing solutions exhibit weak temperature resilience at larger array size and only support 1-bit. In this paper, we propose TReCiM, an ultra-low-power temperature-resilient multibit 2FeFET-1T CiM design that reliably performs MAC operations in the subthreshold-FeFET region with temperature ranging from 0°C to 85°C at scale. We benchmark our design using NeuroSim framework in the context of VGG-8 neural network architecture running the CIFAR-10 dataset. Benchmarking results suggest that when considering temperature drift impact, our proposed TReCiM array achieves 91.31% accuracy, with 1.86% accuracy improvement compared to existing 1-bit 2T-1FeFET CiM array. Furthermore, our proposed design achieves 48.03 TOPS/W energy efficiency at system level, comparable to existing designs with smaller technology feature sizes. Thomas Kämpfe, Kai Ni 0004, Hussam Amrouch, Cheng Zhuo, Xunzhao Yin |
ICCAD | 3 |
| 2024 | Defect Analysis for FeFETs using a Compact ModelabstractFerroelectric field-effect transistors (FeFETs) are a promising emerging non-volatile memory, but the impact of manufacturing imperfections on these devices has yet to be studied comprehensively. We extend a previous FeFET compact model to combine the Preisach ferroelectric capacitor model with the BSIM-SOI MOSFET model. We calibrate this new compact model with data from a technology CAD (TCAD) model that is calibrated against a fabricated metal-ferroelectric-metal capacitor. We analyze polarization defects in the ferroelectric layer using this compact model. We address two classes of defects and map them to stuck-at-fault models, referred to as neutral faults (SAP0), and stuck-at-plus and stuck-at-minus (SAP+and SAP−) faults. This framework obviates the need for computationally expensive TCAD simulations for each defect scenario. Dhruv Thapar, Arjun Chaudhuri, Kai Ni 0004, Krishnendu Chakrabarty |
ITC | 3 |
| 2024 | ProtFe: Low-Cost Secure Power Side-Channel Protection for General and Custom FeFET-Based MemoriesabstractFerroelectric Field Effect Transistors (FeFETs) have spurred increasing interest in both memories and computing applications, thanks to their CMOS compatibility, low-power operation, and high scalability. However, new security threats to the FeFET-based memories also arise. A major threat is the power analysis side-channel attack (P-SCA), which exploits the power traces of the memory access to obtain data information. There have been several effective efforts on resistive nonvolatile memories (NVMs), but they fail to meet the requirements for secure FeFET-based memories due to the different capacitive FeFETs load. Directly applying these existing countermeasures to the P-SCA protection for FeFETs induces huge challenges, especially for the balance between power side-channel resistance and corresponding overheads. To address this issue, we leverage the unique features of FeFETs and propose ProtFe , namely the protection methods for FeFET-based memories, including the pipelined multi-step write strategy ( PiMWrite ) and the split array design ( SpA ). PiMWrite is proposed for general FeFET-based memories, and inserts specially designed intermediate states to mitigate information leakage with pipelined steps to reduce overheads. SpA is proposed for custom FeFET-based memories, and simultaneously writes two split portions of the array with shared minimized peripherals to go beyond the balance between security and overheads. Simulation results show that PiMWrite expands the search space of a single power trace to 21× and involves nearly zero hardware penalties. SpA presents 33× search space improvement with negligible latency, 0.6% area, and only 7.1% energy overhead. ProtFe achieves improved balance between security and overheads, compared with the state-of-the-art works. Taixin Li, Boran Sun, Hongtao Zhong, Yixin Xu 0001, Narayanan Vijaykrishnan, Liang Shi 0001, Thomas Kämpfe, Kai Ni 0004, Huazhong Yang, Xueqing Li 0002 |
ACM Trans. Design Autom. Electr. Syst. | 10 |
| 2024 | A Module-Level Configuration Methodology for Programmable Camouflaged LogicabstractLogic camouflage is a widely adopted technique that mitigates the threat of intellectual property (IP) piracy and overproduction in the integrated circuit (IC) supply chain. Camouflaged logic achieves functional obfuscation through physical-level ambiguity and post-manufacturing programmability. However, discussions on programmability are confined to the level of logic cells/gates, limiting the broader-scale application of logic camouflage. In this work, we propose a novel module-level configuration methodology for programmable camouflaged logic that can be implemented without additional hardware ports and with negligible resources. We prove theoretically that the configuration of the programmable camouflaged logic cells can be achieved through the inputs and netlist of the original module. Further, we propose a novel lightweight ferroelectric FET (FeFET)-based reconfigurable logic gate (rGate) family and apply it to the proposed methodology. With the flexible replacement and the proposed configuration-aware conversion algorithm, this work is characterized by the input-only programming scheme as well as the combination of high output error rate and point-function-like defense. Evaluations show an average of >95% of the alternative rGate location for camouflage, which is sufficient for the security-aware design. We illustrate the exponential complexity in function state traversal and the enhanced defense capability of locked blackbox against Boolean Satisfiability (SAT) attacks compared with key-based methods. We also preserve an evident output Hamming distance and introduce negligible hardware overheads in both gate-level and module-level evaluations under typical benchmarks. Zhonghao Chen, Yixin Xu 0001, Tongguang Yu, Ziheng Zheng, Enze Ye, Sumitha George, Huazhong Yang, Yongpan Liu, Kai Ni 0004, Narayanan Vijaykrishnan, Xueqing Li 0002 |
ACM Trans. Design Autom. Electr. Syst. | 11 |
| 2024 | Multibit Content Addressable Memory Design and Optimization Based on 3-D nand-Compatible IGZO FlashabstractContent addressable memory (CAM) has been employed in various data-intensive tasks for its parallel pattern-matching capability. To enhance the density and efficiency of CAMs, emerging nonvolatile memory (NVM) technologies have been exploited in the CAM designs. Recently, the multilevel cell (MLC) characteristics of NVMs have been utilized in several analog and multibit CAM designs, achieving higher density than conventional binary/ternary CAM designs. However, these analog and multibit CAM designs are built with the practical experience of circuit designers, lacking a general analog/multibit design methodology. In this article, we propose a general and effective design and optimization scheme for multibit CAM, using a novel 3-D nand-compatible amorphous indium–gallium–zinc–oxide (IGZO) flash as a proxy of three-terminal NVM devices. The proposed scheme encodes the multibit data into the flash devices, enabling the 3-D nand flash array to operate as an ultradense nand or nor CAM without significant structural change. For further performance optimization, we propose a design space exploration scheme for optimal CAM parameters. Evaluation results suggest that the CAM design based on our proposed design and optimization scheme achieves over 35$\times$area per bit saving compared with the representative ferroelectric field effect transistor (FeFET)-based multibit CAM, and a 38.1$\times$energy-delay-area product (EDAP) improvement over the state-of-the-art analog CAM, respectively. Chao Li 0065, Chen Sun 0010, Jianyi Yang 0003, Kai Ni 0004, Cheng Zhuo, Xunzhao Yin |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2023 | Cryogenic In-Memory Matrix-Vector Multiplication using Ferroelectric Superconducting Quantum Interference Device (FE-SQUID)abstractNext-generation quantum computing (QC) systems, comprising thousands of qubits, are envisioned to accommodate the quantum substrate (qubits) and classical components (control processor, and a digital memory block) in a cryogenic (< 4 Kelvin) environment. Such homogeneous integration will pave the way for superconducting interconnects and reduce the noise arising from thermal gradient. However, in the existing QC systems, cryogenic control processors and memory blocks are still operated following the von Neumann architecture. This leads to significant performance overhead due to the repetitive data movement between physically distinct memory and processing units. Thus, it becomes challenging to implement computationally expensive machine learning (ML) algorithms for efficient error correction and control of qubits in a QC. In-memory implementation of ML algorithms at cryogenic temperature can be a game-changer for a practical QC. Here, we demonstrate a unique technique for cryogenic in-memory matrix vector multiplication (MVM), the most frequently performed operation in ML algorithms, utilizing a ferroelectric superconducting quantum interference device (FE-SQUID)-based memory array. FE-SQUID is a promising cryogenic memory device thanks to its non-volatile nature, voltage-controlled switching, scalability, and compatibility with commercially available superconducting device fabrication processes. Moreover, due to having separate read-write paths, the read operation can be optimized without imposing any limit on the read bias and hence, multiple levels of read current with notable separation can be used to map the inputs for the MVM operation. We use an experimentally-calibrated compact model for FE-SQUID to design and test our proposed system. We evaluate FE-SQUID-based in-memory MVM by performing several classification tasks using MNIST handwritten digits, fashion, and emotion datasets. We achieve 93.83%, 80.49%, and 92.5% accuracy for handwritten digits, fashion, and sentiment classifications, respectively. Shamiul Alam, Jack Hutchins, Md. Shafayat Hossain, Kai Ni 0004, Narayanan Vijaykrishnan, Ahmedullah Aziz |
DAC | 4 |
| 2023 | SEE-MCAM: Scalable Multi-Bit FeFET Content Addressable Memories for Energy Efficient Associative SearchabstractArtificial intelligence has made remarkable advancements in recent years, leading to the development of algorithms and models capable of handling ever-increasing amounts of data. The computational demands of these algorithms necessitate circuit and architecture designs that go beyond the von-Neumann paradigm. Content addressable memories (CAMs), which implement parallel associative search functionality within memory blocks to overcome the memory wall bottleneck, have proven to be effective for data-intensive tasks. While current CAM designs have achieved higher storage density and energy efficiency than their CMOS-based counterparts by leveraging emerging non-volatile memories (NVM), most of these implementations are limited to binary storage cells. In this work, we propose SEE-MCAM, scalable and compact multi-bit CAM (MCAM) designs that utilize the three-terminal ferroelectric FET (FeFET) as the proxy. By exploiting the multi-level-cell characteristics of FeFETs, our proposed SEE-MCAM designs enable multi-bit associative search functions and achieve better energy efficiency and performance than existing FeFET-based CAM designs. We validated the functionality of our proposed designs by achieving 3 bits per cell CAM functionality, resulting in 3x improvement in storage density. The area per bit of the proposed SEE-MCAM cell is 8% of the conventional CMOS CAM. We thoroughly investigated the scalability and robustness of the proposed design. Evaluation results suggest that the proposed 2FeFET-1 T SEE-MCAM achieves 9.8× more energy efficiency and 1.6× less search latency compared to the CMOS CAM, respectively. When compared to existing MCAM designs, the proposed SEE-MCAM can achieve 8.7× and 4.9× more energy efficiency than ReRAM-based and FeFET-based MCAMs, respectively. Benchmarking results show that our approach provides up to 3 orders of magnitude improvement in speedup and energy efficiency over a GPU implementation in accelerating a novel quantized hyperdimensional computing (HDC) application. Shengxi Shou, Che-Kai Liu, Sanggeon Yun, Zishen Wan, Kai Ni 0004, Mohsen Imani, Xiaobo Sharon Hu, Jianyi Yang 0003, Cheng Zhuo, Xunzhao Yin |
ICCAD | 5 |
| 2023 | FeFET-Based In-Memory Hyperdimensional Encoding DesignabstractThe data explosion of Internet of Things (IoT) and machine learning tasks raises a great demand on highly efficient computing hardware and paradigms. Brain-inspired hyperdimensional computing (HDC) is becoming a promising computing paradigm, which encodes data as hypervectors with homogeneous elements instead of numbers, and can perform learning/classification tasks through simple logical or arithmetic operations on the encoded hypervectors. Therefore HDC has much lower computational complexity than conventional computational models such as neural networks. However, due to its high-dimensional data representation, processing, and encoding hypervectors in conventional Von–Neumann architectures (e.g., CPU and GPU) requires a large amount of energy- and time-consuming data transfer, thus weakening its efficiency benefiting from low complexity. In this article, we proposed an ultralow power and fast computing-in-memory (CiM) design based on nonvolatile (NV) ferroelectric FET (FeFET) for HDC encoding. The proposed design mainly support hyperdimensional bit-wise XOR and parallel majority vote (MAJ) operations for HDC encoding, which are implemented by FeFET-based memories together with CMOS peripheral circuits. The 1FeFET1T-based memory cell effectively mitigates the impact of transistor variations on the operation. A highly parallel and pipelined computing workflow of the proposed design further boosts the energy efficiency and performance with a negligible extra area overhead. Experimental results demonstrate that our proposed design achieves$5.04\times $energy efficiency improvement over other CiM designs for HDC encoding. Qingrong Huang, Kai Ni 0004, Mohsen Imani, Cheng Zhuo, Xunzhao Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Ferroelectric Ternary Content Addressable Memories for Energy-Efficient Associative SearchabstractA fast and efficient search function across the database has been a core component for a number of data-intensive tasks in machine learning, IoT applications, and inference. However, the conventional digital machines implementing the search functionality with repetitive arithmetic operations suffer from the energy efficiency and performance degradation due to the significant data transfer between the storage and processing units in the Von Neumann architecture. Ternary content addressable memories (TCAMs) are an essential hardware form of computing-in-memory (CiM) designs that aim to overcome the data transfer bottlenecks by implementing the parallel associative search function within the memory blocks. While most state-of-the-art TCAM designs focus on improving the information density by harnessing compact nonvolatile memories (NVMs), little efforts have been spent on optimizing the energy efficiency of the NVM-based TCAM. In this article, by exploiting the ferroelectric FET (FeFET) as a representative NVM, we propose an NOR-type 2FeFET-1T and an NAND-type 2FeFET-2T TCAM designs that enable highly energy-efficient associative search by reducing the associated precharge overheads. We then propose a hybrid ferroelectric NAND-NOR (HFNN) TCAM design to further improve the energy efficiency. An HFNN-based segmented architecture is proposed to reduce the search delay and energy by search operation pipeline. Evaluation results suggest that the proposed 2FeFET-1T, 2FeFET-2T and HFNN TCAM design consume$3.03\times $,$8.08\times $, and$226.92\times $less search energy than the conventional 16T complementary metal oxide semiconductor (CMOS) TCAM, respectively. Application benchmarking shows that our proposed 2FeFET-1T/2FeFET-2T/HFNN TCAM can save, on average, 45.2%/50.6%/57.5% the GPU energy consumption as compared to the conventional GPU. Xunzhao Yin, Yu Qian 0002, Mohsen Imani, Kai Ni 0004, Chao Li 0065, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2023 | Design of Ultracompact Content Addressable Memory Exploiting 1T-1MTJ CellabstractContent addressable memories (CAMs) are a promising category of computing-in-memory (CiM) elements that can perform highly parallel and efficient search operations for routers, pattern matching, and other data-intensive applications. Various magnetic tunnel junction (MTJ)-based CAM designs have been proposed to realize zero standby power and high-performance search. However, due to the relatively small tunnel magneto-resistance (TMR) ratio, MTJ-based CAMs require extra transistors and differential MTJ branches to distinguish between the parallel and anti-parallel resistance states, resulting in significant area and energy overhead. In this article, we propose a device-circuit co-design approach for an ultracompact CAM design by only exploiting a 1T-1MTJ structure in each cell. We propose a 2-step search scheme to enable the parallel in-memory search operation across the proposed CAM array and demonstrate the sufficient sensing margin of the array in a successful search operation. Evaluation results suggest that our proposed 1T-1MTJ-based CAM design improves$179\times /301\times $area efficiency compared with the state-of-the-art 15T-4MTJ/20T-6MTJ CAM design. Application benchmarking on hyperdimensional computing (HDC) inference shows a$54.6\times /12.8\times $speedup compared with GPU/20T-6MTJ CAM-based approaches. Cheng Zhuo, Kai Ni 0004, Mohsen Imani, Yuxuan Luo 0001, Shaodi Wang, Deming Zhang, Xunzhao Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2023 | Low-Power and Scalable BEOL-Compatible IGZO TFT eDRAM-Based Charge-Domain ComputingabstractThe rapid development of edge artificial intelligence (AI) raises high requirements for data-intensive neural network (NN) computing and storage of edge devices, under a limited chip footprint and energy supply source. As a promising approach for energy-efficient processing, computing-in-memory (CiM) has been widely explored in recent efforts to mitigate the data transmission bottleneck. However, CiM with small on-chip memory capacity results in expensive data reloads, limiting its deployment in large-scale NN applications. Moreover, the increased leakage under advanced CMOS scaling lowers the energy efficiency. In this work, device-circuit synergy based on the indium-gallium-zinc-oxide (IGZO) thin-film transistor (TFT) is adopted to address these challenges. First, 4-transistor-1-capacitor (4T1C) IGZO eDRAM CiM is proposed with higher density than SRAM-based CiM and enhanced data retention by both lower device leakage and a differential cell structure. Second, exploiting the back-end-of-line (BEOL) compatibility and vertical integration of emerging channel-all-around (CAA) IGZO devices, 3D eDRAM CiM is proposed, which paves the way for IGZO-based CiM with ultra-high density. Circuit techniques including time-interleaved computing and differential refresh are proposed to guarantee accuracy under large-capacity 3D CiM. As a proof of concept, a$128 \times 32$CiM array is fabricated under a foundry low-temperature poly-crystalline and oxide (LTPO) technology, demonstrating high computing linearity and long data retention. Benchmarks on scaled 45nm IGZO technology show energy efficiency of 686 TOPS/W for array only, and 138 TOPS/W while considering peripheral overheads. Jialong Liu, Chen Sun 0010, Yongpan Liu, Huazhong Yang, Kai Ni 0004, Xueqing Li 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 8 |
| 2023 | FeFET-Based Logic-in-Memory Supporting SA-Free Write-Back and Fully Dynamic Access With Reduced Bitline Charging Activity and Recycled Bitline ChargeabstractBitwise logic-in-memory (BLiM) is a promising approach to efficient computing in data-intensive applications by reducing data movement between memory and processing units. However, existing BLiM techniques have challenges towards higher energy efficiency and speed: (i) DC power in computing and result sensing is significant in most existing RRAM and MRAM based BLiM solutions; (ii) before the computation result could be stored back to the same memory array, existing BLiM has to sense the result first, at the cost of extra power and latency due to the sense amplifiers (SAs). Targeting at higher energy efficiency and speed, this work proposes a new BLiM approach in 2-transistor/ cell (2T/C) and 3T/C topologies based on ferroelectric field-effect transistors (FeFETs), supporting a variety of computing functions. For the first time, this new approach supports SA-free direct write-back, and consumes no static power for computing and sensing with proposed fully dynamic computing and sensing schemes. Another highlight is that this work further minimizes the dynamic power by (i) reducing the chance of bitline charging activities and (ii) recycling the bitline charge in sensing multi-operand operations. Compared with prior BLiM methods based on nonvolatile memories, evaluation shows 3.0x–100x latency and 1.3x–200x energy improvement for typical in- memory XOR operation, which further leads to 3.0x–58x and 3.2x–78x savings of latency and energy, respectively, for the application of advanced-encryption standard (AES). Mingyen Lee, Juejian Wu, Yixin Xu 0001, Yongpan Liu, Kai Ni 0004, Yu Wang 0002, Huazhong Yang, Narayanan Vijaykrishnan, Xueqing Li 0002 |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2022 | Energy efficient data search design and optimization based on a compact ferroelectric FET content addressable memoryabstractContent Addressable Memory (CAM) is widely used for associative search tasks in advanced machine learning models and data-intensive applications due to the highly parallel pattern matching capability. Most state-of-the-art CAM designs focus on reducing the CAM cell area by exploiting the nonvolatile memories (NVMs). There exists only little research on optimizing the design and energy efficiency of NVM based CAMs for practical deployment in edge devices and AI hardware. In this paper, we propose a general compact and energy efficient CAM design scheme that alleviates the design overhead by employing just one NVM device in the cell. We also propose an adaptive matchline (ML) precharge and discharge scheme that further optimizes the search energy by fully reducing the ML voltage swing. We consider Ferroelectric field effect transistors (FeFETs) as the representative NVM, and present a 2T-1FeFET CAM array including a sense amplifier implementing the proposed ML scheme. Evaluation results suggest that our proposed 2T-1FeFET CAM design achieves 6.64×/4.74×/9.14×/3.02× better energy efficiency compared with CMOS/ReRAM/STT-MRAM/2FeFET CAM arrays. Benchmarking results show that our approach provides 3.3×/2.1× energy-delay product improvement over the 2T-2R/2FeFET CAM in accelerating query processing applications. Jiahao Cai, Mohsen Imani, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin |
DAC | 3 |
| 2022 | Adaptive neural recovery for highly robust brain-like representationabstractToday's machine learning platforms have major robustness issues dealing with insecure and unreliable memory systems. In conventional data representation, bit flips due to noise or attack can cause value explosion, which leads to incorrect learning prediction. In this paper, we propose RobustHD, a robust and noise-tolerant learning system based on HyperDimensional Computing (HDC), mimicking important brain functionalities. Unlike traditional binary representation, RobustHD exploits a redundant and holographic representation, ensuring all bits have the same impact on the computation. RobustHD also proposes a runtime framework that adaptively identifies and regenerates the faulty dimensions in an unsupervised way. Our solution not only provides security against possible bit-flip attacks but also provides a learning solution with high robustness to noises in the memory. We performed a cross-stacked evaluation from a conventional platform to emerging processing in-memory architecture. Our evaluation shows that under 10% random bit flip attack, RobustHD provides a maximum of 0.53% quality loss, while deep learning solutions are losing over 26.2% accuracy. Prathyush Poduval, Yang Ni 0001, Yeseong Kim, Kai Ni 0004, Raghavan Kumar, Rosario Cammarota, Mohsen Imani |
DAC | 4 |
| 2022 | Eva-CAM: A Circuit/Architecture-Level Evaluation Tool for General Content Addressable MemoriesabstractContent addressable memories (CAMs), a special-purpose in-memory computing (IMC) unit, support parallel searches directly in memory. There are growing interests in CAMs for data-intensive applications such as machine learning and bioinformatics. The design space for CAMs is rapidly expanding. In addition to traditional ternary CAMs (TCAMs), analog CAM (ACAM) and multi-bit CAM (MCAM) designs based on various non-volatile memory (NVM) devices have been recently introduced and may offer higher density, better energy efficiency, and non-volatility. Furthermore, aside from the widely-used exact match based search, CAM-based approximate matches have been proposed to further extend the utility of CAMs to new application spaces. For this memory architecture, evaluating different CAM design options for a given application is becoming more challenging. This paper presents Eva-CAM, a circuit/architecture-level modeling and evaluation tool for CAMs. Eva-CAM supports TCAM, ACAM, and MCAM designs implemented in non-volatile memories, for both exact and approximate match types. It also allows for the exploration of CAM array structures and sensing circuits. Eva-CAM has been validated with HSPICE simulation results and chip measurements. A comprehensive case study is described for FeFET CAM design space exploration. Liu Liu 0023, Mohammad Mehdi Sharifi, Ramin Rajaei, Arman Kazemi, Kai Ni 0004, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 5 |
| 2022 | COSIME: FeFET Based Associative Memory for In-Memory Cosine Similarity SearchabstractIn a number of machine learning models, an input query is searched across the trained class vectors to find the closest feature class vector in cosine similarity metric. However, performing the cosine similarities between the vectors in Von-Neumann machines involves a large number of multiplications, Euclidean normalizations and division operations, thus incurring heavy hardware energy and latency overheads. Moreover, due to the memory wall problem that presents in the conventional architecture, frequent cosine similarity-based searches (CSSs) over the class vectors requires a lot of data movements, limiting the throughput and efficiency of the system. To overcome the aforementioned challenges, this paper introduces COSIME, a general in-memory associative memory (AM) engine based on the ferroelectric FET (FeFET) device for efficient CSS. By leveraging the one-transistor AND gate function of FeFET devices, current-based translinear analog circuit and winner-take-all (WTA) circuitry, COSIME can realize parallel in-memory CSS across all the entries in a memory block, and output the closest word to the input query in cosine similarity metric. Evaluation results at the array level suggest that the proposed COSIME design achieves 333× and 90.5× latency and energy improvements, respectively, and realizes better classification accuracy when compared with an AM design implementing approximated CSS. The proposed in-memory computing fabric is evaluated for an HDC problem, showcasing that COSIME can achieve on average 47.1× and 98.5× speedup and energy efficiency improvements compared with an GPU implementation. Che-Kai Liu, Haobang Chen, Mohsen Imani, Kai Ni 0004, Arman Kazemi, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu, Liang Zhao 0004, Cheng Zhuo, Xunzhao Yin |
ICCAD | 4 |
| 2022 | Carbon Nanotube SRAM in 5-nm Technology Node Design, Optimization, and Performance Evaluation - Part I: CNFET Transistor OptimizationabstractIn this article, we propose a carbon nanotube (CNT) field-effect transistor (CNFET)-based static random access memory (SRAM) design at the 5-nm technology node that is optimized based on the tradeoff between performance, stability, and power efficiency. In addition to size optimization, physical model parameters including CNT density, CNT diameter, and CNFET flat band voltage are evaluated and optimized for CNFET SRAM performance improvement. Optimized CNFET SRAM is compared with state-of-the-art 7-nm FinFET SRAM cell based on Arizona State University [ASAP 7-nm FinFET predictive technology models (PTM)] library. We find that the read, write EDPs, and static power of the proposed CNFET SRAM cell are improved by 67.6%, 71.5%, and 43.6%, respectively, compared with the FinFET SRAM cell, with slightly better stability. CNT interconnects both inside and in-between CNFET SRAM cells are considered to compose an all-carbon-based SRAM (ACS) array which will be discussed in the Part II of this article. A 7-nm FinFET SRAM cell with copper interconnects is implemented and used for comparison. Rongmei Chen, Yuanqing Cheng, Souhir Elloumi, Kangwei Xu, Vihar P. Georgiev, Kai Ni 0004, Peter Debacker, A. Asenov, Aida Todri |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2022 | Carbon Nanotube SRAM in 5-nm Technology Node Design, Optimization, and Performance Evaluation - Part II: CNT Interconnect OptimizationabstractThe size and parameter optimization for the 5-nm carbon nanotube field effect transistor (CNFET) static random access memory (SRAM) cell was presented in Part I of this article. Based on that work, we propose a carbon nanotube (CNT) SRAM array composed of the schematically optimized CNFET SRAM and CNT interconnects. We consider the interconnects inside the CNFET SRAM cell composed of metallic single-wall CNT (M-SWCNT) bundles to represent the metal layers 0 and 1 (M0 and M1). We investigate the layout structure of CNFET SRAM cell considering CNFET devices, M-SWCNT interconnects, and metal electrode Palladium with CNT (Pd-CNT) contacts. Two versions of cell layout designs are explored and compared in terms of performance, stability, and power efficiency. Furthermore, we implement a 16 Kbit SRAM array composed of the proposed CNFET SRAM cells, multiwall CNT (MWCNTs) inter-cell interconnects and Pd-CNT contacts. Such an array shows significant advantages, with the read and write overall energy-delay product (EDP), static power consumption, and core area of$0.28\times $,$0.52\times $, and$0.76\times $respectively to 7-nm FinFET-SRAM array with copper interconnects, whereas the read and write static noise margins are 6% and 12% respectively larger than the FinFET counterpart. Rongmei Chen, Yuanqing Cheng, Souhir Elloumi, Kangwei Xu, Vihar P. Georgiev, Kai Ni 0004, Peter Debacker, A. Asenov, Aida Todri |
IEEE Trans. Very Large Scale Integr. Syst. | 9 |
| 2022 | Machine Learning Attack Resistant Area-Efficient Reconfigurable Ising-PUFabstractThe Ising-physical unclonable function (PUF) is a recent PUF structure formed of a network of APUFs inspired by the Ising model. A large challenge–response pair (CRP) space with high resilience against machine learning modeling attacks can be attained due to the unique arrangement. These advantages, however, are achieved at the cost of a large area overhead. In this article, a reconfigurable Ising-PUF is introduced with several new design knobs to generate a much larger CRP space within a smaller area. A 25% increase in the number of challenge bits can be achieved for a design that occupies 30% of the conventional Ising-PUF area. With the proposed lightweight area-efficient design, up to 5.6 times lower area per CRP can be achieved compared to the existing design. Several improvements are proposed that leverage the large design space, enabling dynamic tradeoffs between the area and CRP pool with the proposed flexible customization of Ising-PUFs. A detailed analysis of this improved design space is explored for different parameters. The state-of-the-art machine learning modeling attacks are investigated, and the Ising-PUF structure is shown to be resilient. Eslam Elmitwalli, Kai Ni 0004, Selçuk Köse |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | CapCAM: A Multilevel Capacitive Content Addressable Memory for High-Accuracy and High-Scalability Search and Compute ApplicationsabstractAs one type of associative memory, content-addressable memory (CAM) has become a critical component in several applications, including caches, routers, and pattern matching. Compared with the conventional CAM that could only deliver a “matched or not-matched” result, emerging multilevel CAM (ML-CAM) is capable of delivering “the degree of match” with multilevel distance calculation. This feature has been desired in applications that need beyond-Boolean matching results. However, existing ML-CAM designs are limited by the bit-cell device discharging current mismatch and vulnerability to the timing of sensing operations for distance calculation. This inherent constraint makes it difficult to further improve the accuracy and scalability toward higher accuracy and higher dimension matching. In this work, we propose CapCAM, a multilevel Capacitive Content Addressable Memory. It could be implemented based on either static random-access memory (SRAM) or emerging technologies, e.g., the ferroelectric field-effect transistor (FeFET). CapCAM could provide linear and stable voltage drop scaled by the match degree and need no strict timing for result sensing, which embraces the high-accuracy and high-scalability search. The inherent enabler of CapCAM is the charge-domain computing mechanism. This article will present the basic concept, operating mechanisms, detailed circuit designs, and circuit-level simulations of CapCAM. Besides, we apply CapCAM to few-shot learning applications and compare CapCAM with the current-domain TCAM designs. Results show 99.2% accuracy for a five-way five-shot classification task with our proposed CapCAM design while considering 1-fF capacitors, 20-domain FeFETs, and 256 columns. In contrast, the prior work based on discharging dynamics requires strict timing controls and suffers from accuracy degradation under the same configuration, which demonstrates CapCAM’s capability of low-power, accurate, and scalable multilevel CAM (ML-CAM) computing. Hongtao Zhong, Nuo Xiu, Guodong Yin, Narayanan Vijaykrishnan, Yongpan Liu, Kai Ni 0004, Huazhong Yang, Xueqing Li 0002 |
IEEE Trans. Very Large Scale Integr. Syst. | 8 |
| 2021 | Energy-Aware Designs of Ferroelectric Ternary Content Addressable MemoryabstractTernary content addressable memories (TCAMs) are a special form of computing-in-memory (CiM) circuits that aim to address the so-called memory wall issues by merging the parallel search function with memory blocks. Due to the content addressing nature, TCAMs have been widely utilized for search intensive tasks in low-power, data analytic applications, such as IP routers, associative memories, and learning models. While most state-of-the-art TCAM designs focus on improving the TCAM density by harnessing compact nonvolatile memories (NVMs), little efforts have been spent on reducing and optimizing the energy consumption of the NVM based TCAM. In this paper, by exploiting the Ferroelectric FET (FeFET) as a representative NVM, we propose two compact and energy-aware designs of ferroelectric TCAMs for low power applications. We first introduce a novel 2FeFET based XOR-like gate structure that can also be adopted to other NVMs, and then leverage the structure to propose two TCAM designs that achieve high energy efficiency by either reducing the associated precharge overhead (2FeFET-1T cell), or eliminating the precharge phase typically required by TCAMs (2FeFET-2T cell). We evaluate and compare the designs w.r.t area, search energy and delay at array level with other existing designs, and benchmark the proposed TCAM designs in an associative memory based GPU architecture. The results suggest that the proposed 2FeFET-1T/2FeFET-2T TCAM design consumes 3.03X/8.08X less search energy than the conventional 16T CMOS TCAM, while the proposed design cell area is only 32.1%/39.3% of the latter. Compared with the state-of-the-art 2FeFET only TCAM array, our proposed designs still achieve 1.79X and 4.79X search energy reduction, respectively. Moreover, our proposed designs can achieve, on average, 45.2%/51.5% energy saving compared with the conventional GPU based architecture at the application level. Yu Qian 0002, Zhenhao Fan, Chao Li 0065, Mohsen Imani, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin |
DATE | 6 |
| 2021 | Overview of Ferroelectric Memory Devices and Reliability Aware Design OptimizationabstractSince the discovery of CMOS-compatible and highly scalable ferroelectric HfO2, there has been a significant revival of interest in developing ferroelectric devices for high performance and energy-efficient embedded nonvolatile memories. Multiple ferroelectric memory devices are under investigation by harnessing the nonvolatile polarization states. These devices include the ferroelectric FET (FeFET), ferroelectric capacitor based random access memory (FeRAM), and ferroelectric tunnel junction (FTJ). Though the underlying memory storage mechanisms are the same in these devices, their memory sensing mechanisms are different. This difference leads to fundamentally different, and even opposite ferroelectric optimization directions. Given their different characteristics and individual advantages, it is likely that all these devices will co-exist to meet varying needs. Therefore, it is important to establish and compare the design guidelines for the three different ferroelectric memory devices. The design optimization will also be constrained by the reliability limit, which is critical to guarantee the success of the ferroelectric memory devices. Shan Deng, Zijian Zhao 0001, Santosh Kurinec, Kai Ni 0004, Yi Xiao 0008, Tongguang Yu, Narayanan Vijaykrishnan |
ACM Great Lakes Symposium on VLSI | 4 |
| 2021 | ICCAD Tutorial Session Paper Ferroelectric FET Technology and Applications: From Devices to SystemsabstractThe rapidly increasing volume and complexity of data is demanding the relentless scaling of computing power. With transistor feature size approaching physical limits, the benefits that CMOS technology can provide is diminishing. For future energy efficient computing systems, researchers aim to exploit various emerging nanotechnologies to replace conventional CMOS technology. In particular, ferroelectric FETs (FeFETs) appear to be a promising candidate to continue improving energy efficiency for data-intensive applications. Advances in FeFET scalability and FeFET compatibility with CMOS have sparked growing interest in device, circuit, and system communities. While FeFET is still evolving, many researchers and developers are already cautiously optimistic about its future. This paper provides a review on FeFET's recent technology advances, challenges, and opportunities, with a particular emphasis upon device modeling and circuit design of FeFET content addressable memory, as well as their applications in machine learning. Hussam Amrouch, Xiaobo Sharon Hu, Arman Kazemi, Ann Franchesca Laguna, Kai Ni 0004, Michael T. Niemier, Mohammad Mehdi Sharifi, Simon Thomann, Xunzhao Yin, Cheng Zhuo |
ICCAD | 6 |
| 2021 | Application-driven Design Exploration for Dense Ferroelectric Embedded Non-volatile MemoriesabstractThe memory wall bottleneck is a key challenge across many data-intensive applications. Multi-level FeFET-based embedded non-volatile memories are a promising solution for denser and more energy-efficient on-chip memory. However, reliable multi-level cell storage requires careful optimizations to minimize the design overhead costs. In this work, we investigate the interplay between FeFET device characteristics, programming schemes, and memory array architecture, and explore different design choices to optimize performance, energy, area, and accuracy metrics for critical data-intensive workloads. From our cross-stack design exploration, we find that we can store DNN weights and social network graphs at a density of over 8MB/mm2and sub-2ns read access latency without loss in application accuracy. Mohammad Mehdi Sharifi, Lillian Pentecost, Ramin Rajaei, Arman Kazemi, Qiuwen Lou, Gu-Yeon Wei, David Brooks 0001, Kai Ni 0004, Xiaobo Sharon Hu, Michael T. Niemier, Marco Donato |
ISLPED | 8 |
| 2020 | Ferroelectrics: From Memory to ComputingabstractResearch discovery of ferroelectricity in doped hafnium dioxide thin films has ignited tremendous activity in exploration of ferroelectric FETs for a range of applications from low-power logic to embedded non-volatile memory to in-memory compute kernels. In this paper, key milestones in the evolution of Ferroelectric Field Effect Transistors (FeFETs) and the emergence of a versatile ferroelectronic platform are presented. FeFET exhibits superior energy efficiency and high performance as embedded nonvolatile memory. When embedded into logic, such as SRAM or D-flip-flop, nonvolatile processor can be designed, which is critical for intermittent computing with unreliable power. The partial polarization switching in multi-domain ferroelectric can be harnessed to develop analog synaptic weight cell for deep learning accelerators. To further improve the energy-efficiency of computation, ferroelectric in-memory computing hardware primitive is designed, with one prominent example of ferroelectric TCAM. Utilizing the ferroelectric switching dynamics, ferroelectric neuron with intrinsic homeostasis can be realized to enable a unified ferroelectric platform for spiking neural network. From all these developments, ferroelectric emerges as a highly promising platform for various exciting applications. Kai Ni 0004, Suman Datta |
ASP-DAC | 1 |
| 2020 | A Hybrid FeMFET-CMOS Analog Synapse Circuit for Neural Network Training and InferenceabstractAn analog synapse circuit based on ferroelectric-metal field-effect transistors is proposed, that offers 6-bit weight precision. The circuit is comprised of volatile least significant bits (LSBs) used solely during training, and non-volatile most significant bits (MSBs) used for both training and inference. The design works at a 1.8V logic-compatible voltage, provides 1010endurance cycles, and requires only 250ps update pulses. A variant of LeNet trained with the proposed synapse achieves 98.2% accuracy on MNIST, which is only 0.4% lower than an ideal implementation of the same network with the same bit precision. Furthermore, the proposed synapse offers improvements of up to 26% in area, 44.8% in leakage power, 16.7% in LSB update pulse duration, and two orders of magnitude in endurance cycles, when compared to state-of-the-art hybrid synaptic circuits. Our proposed synapse can be extended to an 8-bit design, enabling a VGG-like network to achieve 88.8% accuracy on CIFAR-10 (only 0.8% lower than an ideal implementation of the same network). Arman Kazemi, Ramin Rajaei, Kai Ni 0004, Suman Datta, Michael T. Niemier, Xiaobo Sharon Hu |
ISCAS | 3 |
| 2019 | A 3T/Cell Practical Embedded Nonvolatile Memory Supporting Symmetric Read and Write Access Based on Ferroelectric FETsabstractMaking embedded memory symmetric provides the capability of memory access in both rows and columns, which brings new opportunities of significant energy and time savings if only a portion of data in the words need to be accessed. This work investigates the use of ferroelectric field-effect transistors (FeFETs), an emerging nonvolatile, low-power, deeply-scalable, CMOS-compatible transistor technology, and proposes a new 3-transistor/cell symmetric nonvolatile memory (SymNVM). With ~1.67x higher density as compared with the prior FeFET design, significant benefits of energy and latency improvement have been achieved, as evaluated and discussed in depth in this paper. Juejian Wu, Hongtao Zhong, Kai Ni 0004, Yongpan Liu, Huazhong Yang, Xueqing Li 0002 |
DAC | 3 |
| 2018 | Computing with ferroelectric FETs: Devices, models, systems, and applicationsabstractIn this paper, we consider devices, circuits, and systems comprised of transistors with integrated ferroelectrics. Said structures are actively being considered by various semiconductor manufacturers as they can address a large and unique design space. Transistors with integrated ferroelectrics could (i) enable a better switch (i.e., offer steeper subthreshold swings), (ii) are CMOS compatible, (iii) have multiple operating modes (i.e., I-V characteristics can also enable compact, 1-transistor, non-volatile storage elements, as well as analog synaptic behavior), and (iv) have been experimentally demonstrated (i.e., with respect to all of the aforementioned operating modes). These device-level characteristics offer unique opportunities at the circuit, architectural, and system-level, and are considered here from device, circuit/architecture, and foundry-level perspectives. Ahmedullah Aziz, Evelyn T. Breyer, Xiaoming Chen 0003, Suman Datta, Sumeet Kumar Gupta, Michael Hoffmann 0008, Xiaobo Sharon Hu, Adrian M. Ionescu, Matthew Jerry, Thomas Mikolajick, Halid Mulaosmanovic, Kai Ni 0004, Michael T. Niemier, Ian O'Connor, Atanu Saha, Stefan Slesazeck, Sandeep Krishna Thirumala, Xunzhao Yin |
DATE | 13 |