VLDB 2026 Research / reviewers in the wild / expert
Xunzhao Yin
dblp:179/2993
· DBLP profile ↗
99ranked-venue papers
8as first author
76since 2021 · last 2026
0000-0003-4656-9545ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 97 · 8 first-author · 74 since 2021Software engineering, systems software and programming languages · 18 · 2 first-author · 12 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CDACiM: A Charge-Domain Compute-in-Memory Macro for FP/INT MAC Operations with Reconfigurable Capacitor Digital-Analog-ConverterabstractAdvanced edge artificial intelligence (AI) chips need to balance flexible computation, high energy efficiency, and sufficient inference accuracy across diverse workloads. Many compute-in-memory (CiM) designs enable efficient neural network acceleration but focus solely on integer (INT) multiply-and-accumulate (MAC) operations, limiting precision. Some CiM macros add extra circuitry to support floating point (FP) MACs, but these dedicated exponent-handling blocks often waste area when running INT workloads. In this paper, we propose CDACiM, a charge-domain CiM macro that supports both FP and INT MAC operations with minimal overhead. CDACiM introduces a reconfigurable capacitor digital-to-analog converter (RCDAC) that performs both exponent summation and bitwise AND for mantissa multiplication. To calculate exponent offsets, we develop a shared single-slope ADC (SS-ADC) that finds the maximum exponent and computes differences in time domain simultaneously. Our design includes a sparsity-aware computation scheme with tunable thresholds that skips low-importance input-weight pairs, boosting energy efficiency through higher input sparsity. We also introduce a multi-bit input accumulation method that leverages ADC redundancy during quantization and normalization to improve performance. Implemented in a 40nm CMOS process, CDACiM demonstrates an excellent flexibility and trade-off between accuracy and resource usage. Notably, it is the first CiM design to reconfigure capacitor-based INT macro for parallel exponent computation. CDACiM achieves 16.2 TOPS/W for INT MACs and $\mathbf{1 5. 9}$ TFLOPS/W for FP MACs. It delivers a $\mathbf{1. 3 6 - 1. 4 8 \times}$ improvement in energy efficiency with minimal accuracy loss compared to recent FP CiM macros. Jinting Yao, Yuxiao Jiang, Zheyu Yan, Cheng Zhuo, Xunzhao Yin |
ASP-DAC | 7 |
| 2026 | FeFET-Based Analog In-Memory Computing With Inherent Shift-and-Add CapabilityabstractIn-memory computing (IMC) architecture has emerged as a highly promising approach, enhancing the energy efficiency of multiply-and-accumulate (MAC) operations in deep neural networks (DNNs) by embedding parallel computations directly into memory arrays. However, existing ferroelectric FET (FeFET)-based analog IMC designs are often constrained to cell-level optimizations and struggle to achieve high-precision MAC operations. In contrast, high-precision analog IMC architectures typically perform MAC operations for partial inputs and weights within the array in a single cycle and then accumulate partial results over multiple cycles. During this procedure, circuits that handle weight shift-and-add process, whether in digital or analog form, incur significant overhead. This paper presents energy-efficient high-precision analog IMC designs leveraging FeFET technology, which inherently support a shift-and-add mechanism for weights. Initially, we introduce an IMC array paradigm that performs partial MAC operations within each column, and seamlessly incorporates the shift-and-add process for weights by utilizing the analog storage properties of FeFET-based cells. Building upon this paradigm, we propose single-level cell (SLC) FeFET-based designs, namely CurFe and ChgFe, operating in the current and charge modes, respectively. Additionally, to leverage FeFET’s multi-level cell (MLC) properties, we propose a novel hybrid SLC-MLC FeFET-based design, MulFe, which offers higher storage density and energy efficiency. Comprehensive evaluations are conducted at both the circuit and system levels, and the results indicate that the average energy efficiency of the proposed FeFET-based analog IMC designs is 1.32× to 2.71× higher compared to state-of-the-art (SOTA) IMC designs. Qingrong Huang, Yu Qian 0002, Jiahao Cai, Kai Ni 0004, Thomas Kämpfe, Zheyu Yan, Xunzhao Yin, Cheng Zhuo |
IEEE Trans. Computers | 9 |
| 2026 | EvaCAM: A Circuit-Level Evaluation Tool for General Content Addressable MemoriesabstractContent addressable memories (CAMs) are special-purpose in-memory computing units that support parallel searches directly in memory. There is growing interest in CAMs for data-intensive applications such as machine learning, data mining, and bioinformatics, which has led to a rapidly growing CAM design space. CAM cells can be implemented exclusively by CMOS or with various non-volatile memory (NVM) devices. In addition to traditional binary and ternary CAMs (BCAMs and TCAMs), analog CAM (ACAM) and multi-bit CAM (MCAM) designs have recently been introduced, which could further improve density, and also support unique in-memory distance functions. Furthermore, aside from the widely-used exact match function, CAM-based approximate match functions, such as threshold match and best match, have been proposed to further extend the utility of CAMs to new application spaces. As the CAM design space is large, evaluating different CAM design options for a given application is both crucial and challenging. This paper presents EvaCAM, a circuit-level modeling and evaluation tool for CAMs. EvaCAM supports TCAM, ACAM, and MCAM designs implemented in either CMOS or NVMs, for both exact and approximate match functions. It also allows for the exploration of different CAM designs under various optimization targets. EvaCAM has been validated against measured data from fabricated chips and detailed SPICE simulations. A comprehensive design space exploration for CAMs is provided to illustrate the impact of various design decisions and to demonstrate the use cases of EvaCAM. Liu Liu 0023, Mohammad Mehdi Sharifi, Kunshi Wang, Ruibin Mao, Kai Ni 0004, Can Li 0024, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 7 |
| 2026 | ZlibBoost: An Efficient and Flexible Open-Source Framework for Standard Cell CharacterizationabstractAs VLSI designs grow increasingly complex and transition to smaller process nodes, accurate and efficient library characterization has become essential for modern design workflows. Existing open-source tools are often constrained by limited functionality, efficiency, and accuracy, making them insufficient for today’s design challenges. This article reviews the shortcomings of current open-source tools and introduces ZlibBoost, a novel open-source framework designed to provide both flexibility and high performance. Its modular, front-end and back-end separated architecture, along with user-friendly interfaces, enables seamless customization, integration of machine learning models, and expanded simulator compatibility. A variety of key features are introduced to significantly enhance both accuracy and efficiency of library characterization. Experimental results demonstrate ZlibBoost’s capability to meet the demands of both academic research and practical applications, establishing it as a robust solution for advancing semiconductor design. Zhengrui Chen, Chengjun Guo, Shizhang Wang, Guozhu Feng, Zixuan Song, Xunzhao Yin, Weiquan Song, Li Zhang 0021, Zheyu Yan, Cheng Zhuo |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2026 | Machine Learning-Assisted VCD Processing for Accelerated Dynamic Voltage Drop AnalysisabstractWith escalating power integrity challenges in advanced technologies, acquiring accurate dynamic power supply noise through Dynamic Voltage Drop (DVD) analysis becomes increasingly demanding. As noise margins shrink, the use of Value Change Dump (VCD) files for precise DVD analysis is indispensable but computationally expensive. Furthermore, the substantial storage requirements of VCD files, which record digital waveforms from logical simulations, pose significant challenges. In this article, we propose a machine learning (ML)-assisted VCD processing framework to accelerate DVD analysis and improve data efficiency. Transitions recorded in VCD files are mapped to a Physical Design-Aware Circuit Hierarchy Tree (CHT) for efficient feature extraction. These features are leveraged by an XGBoost-based predictor to identify critical vector time windows within the VCD, significantly reducing simulation complexity. Additionally, Huffman encoding is applied to compress signal names, further optimizing storage utilization. Experimental results show that DVD analysis using our profiled VCD files achieves a speedup of approximately 3.53× with an error margin of only 3.89%. Jingchao Hu, Yufei Chen 0007, Songyu Sun, Jianfei Song, Li Zhang 0021, Xunzhao Yin, Zhou Jin 0001, Cheng Zhuo |
ACM Trans. Design Autom. Electr. Syst. | 6 |
| 2026 | HLSRewriter: Efficient Refactoring and Optimization of C/C++ Code with LLMs for High-Level SynthesisabstractIn High-Level Synthesis (HLS), refactoring a standard C/C++ code into its HLS-compatible version (HLS-C) still requires significant human effort. While various program scripts have been introduced to automate this process, the resulting code still contains many HLS-incompatible issues that need to be manually refactored and optimized by developers. Since Large Language Models (LLMs) have the ability to automate code generation, they can also be used for automated code refactoring and optimization in HLS. However, due to the limited training of LLMs, considering hardware and software simultaneously, hallucinations may occur when using LLMs for HLS, leading to synthesis failures. To address these challenges, we introduce HLSRewriter , an LLM-aided code refactoring and optimization framework that takes regular C/C++ code as input and automatically generates its corresponding optimized HLS-C code for hardware synthesis with minimal human intervention. To mitigate LLM hallucinations, a step-wise reasoning process is employed to analyze and detect HLS-incompatible errors. Afterwards, a repair library containing reference templates is efficiently created by scanning the HLS tool manual, followed by cooperation with a Retrieval-Augmented Generation (RAG) paradigm to guide the LLMs toward correct refactoring. In addition, a pipeline-aware decomposition strategy is introduced to progressively break down complex loop structures into smaller tasks with a balanced trade-off between latency and area, thereby enabling efficient pipelining and parallel execution. To further improve hardware efficiency, a bit width adjuster module is incorporated into this framework to optimize the precision of floating-point variables. Moreover, LLM-aided HLS optimization strategies are introduced to add/tune hardware directives in HLS-C code, thereby enhancing the performance of the final synthesized hardware. Experimental results demonstrate that the proposed LLM-aided framework can achieve higher refactoring pass rates and superior hardware performance in 24 real-world tasks compared with traditional approaches and the direct application of LLMs for code refactoring and optimization. The codes are open-sourced at this link: https://github.com/code-source1/catapult . Kangwei Xu, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Ulf Schlichtmann, Bing Li 0005 |
ACM Trans. Design Autom. Electr. Syst. | 3 |
| 2025 | Invited Paper: Boosting Standard Cell Library Characterization with Machine LearningabstractAs VLSI designs grow more complex and transition to smaller process nodes, accurate and efficient library characterization has become increasingly crucial within DTCO and STCO flows. Current open-source tools, however, are constrained to basic library characterization functions and fail to adequately meet modern design demands. In this paper, we review the existing open-source standard cell characterization tools, summarize their limitations, and introduce ZlibBoost---a new open-source framework designed to offer both flexibility and efficiency. We leverage ZlibBoost for LUT index optimization, dynamic power supply noise modeling, and machine learning-based prediction to enhance efficiency and accuracy in library characterization. Experimental results show that such a tool is helpful for both academia and industry to effectively navigate DTCO and STCO challenges. Zhengrui Chen, Chengjun Guo, Zixuan Song, Guozhu Feng, Shizhang Wang, Li Zhang 0021, Xunzhao Yin, Zheyu Yan, Cheng Zhuo |
ASP-DAC | 7 |
| 2025 | VQT-CiM: Accelerating Vector Quantization Enhanced Transformer with Ferroelectric Compute-in-MemoryabstractTransformer models have achieved state-of-the-art performance in various natural language processing (NLP) and computer vision (CV) tasks. To meet their substantial computational demands, the compute-in-memory (CiM) architectures, which alleviate the memory wall problem and enable efficient vector-matrix multiplication (VMM), have been adopted for transformer accelerators. However, the dynamic VMM involved in the attention mechanism, which necessitates runtime write operations, presents significant challenges for non-volatile memory (NVM)-based CiM designs. High write overhead, complex compute-write-compute (CWC) dependencies, and limited endurance reduce their effectiveness. In this paper, we propose VQT-CiM, a ferroelectric FET (FeFET)-based CiM design that accelerates vector quantization (VQ) enhanced transformers by eliminating the runtime write operations. VQT-CiM quantizes keys and values in self-attention to convert dynamic VMMs in inner-product and weighted-sum into static VMMs with the codebooks, enabling efficient calculations with CiM crossbars. However, directly applying VQ hinders the accuracy of transformer model due to its limited representation capability. To address this, we introduce a vector quantization scheme that integrates residual VQ (RVQ) and product VQ (PVQ) for enhanced representation space. We present an efficient hardware implementation for the proposed VQT-CiM with optimized dataflow in RVQ, which incorporates the FeFET-based CiM crossbars and peripheral digital circuits. Evaluation results suggest that VQT-CiM achieves the $3.54 \times$ and $4.53 \times$ improvements in energy efficiency and throughput, respectively, compared to state-of-the-art NVM-based CiM transformer designs. Xuchu Huang, Haonan Du, Zheyu Yan, Cheng Zhuo, Xunzhao Yin |
DAC | 6 |
| 2025 | Device-Algorithm Co-Design of Ferroelectric Compute-in-Memory In-Situ Annealer for Combinatorial Optimization ProblemsabstractCombinatorial optimization problems (COPs) are crucial in many applications but are computationally demanding. Traditional Ising annealers address COPs by directly converting them into Ising models (known as direct-E transformation) and solving them through iterative annealing. However, these approaches require vector-matrix-vector (VMV) multiplications with a complexity of $O\left(n^{2}\right)$ for Ising energy computation and complex exponential annealing factor calculations during annealing process, thus significantly increasing hardware costs. In this work, we propose a ferroelectric compute-in-memory (CiM) in-situ annealer to overcome aforementioned challenges. The proposed device-algorithm co-design framework consists of (i) a novel transformation method (first to our known) that converts COPs into an innovative incremental-E form, which reduces the complexity of VMV multiplication from $O\left(n^{2}\right)$ to $O(n)$, and approximates exponential annealing factor with a much simplified fractional form; (ii) a double gate ferroelectric FET (DG FeFET)-based CiM crossbar that efficiently computes the in-situ incremental-E form by leveraging the unique structure of DG FeFETs; (iii) a CiM annealer that approaches the solutions of COPs via iterative incremental-E computations within a tunable back gate-based in-situ annealing flow. Evaluation results show that our proposed CiM annealer significantly reduces hardware overhead, reducing energy consumption by $1503 / 1716 \times$ and time cost by $8.08 / 8.15 \times$ in solving 3000 -node Max-Cut problems compared to two state-of-the-art annealers. It also exhibits high solving efficiency, achieving a remarkable average success rate of $98 \%$, whereas other annealers show only $50 \%$ given the same iteration counts. Yu Qian 0002, Xianmin Huang, Thomas Kämpfe, Cheng Zhuo, Xunzhao Yin |
DAC | 8 |
| 2025 | FeKAN: Efficient Kolmogorov-Arnold Networks Accelerator Using FeFET-based CAM and LUTabstractKolmogorov-Arnold networks (KANs) have emerged as a promising alternative to MLP due to their adaptive learning capabilities for complex dependencies through B-spline basis activations (BBA). However, existing in-memory accelerators optimized for MLP-based DNNs are primarily designed for vector-matrix multiplication (VMM), making them inefficient for the dynamic and recursive B-spline interpolation (BSI) operations required by KANs. In this work, we propose FeKAN, an FeFET-based architecture designed to accelerate BBA operations. First, we develop a software-hardware co-optimized framework for mapping B-spline basis functions (BBF), leveraging a two-stage design space exploration (DSE) algorithm in combination with FeFET-based Look-Up Tables (LUT) and Content-Addressable Memory (CAM). This framework translated dynamic BSI operations into static codebook lookups, achieving a balanced trade-off between memory and computational efficiency. Second, we propose compress-sparsity-column (CSC) based encoding for B-spline basis function and grouped-computation strategy for memory and energy reduction. Third, we propose a groupedpipeline optimization strategy to mitigate data dependencies, significantly enhancing computation efficiency. Experimental results demonstrate that FeKAN achieves up to $150.68 \mathrm{~K} \times$ and $4664 \times$ higher throughput and up to $606.87 \times$ and $11196 \times$ greater energy efficiency over Intel Xeon Silver 4310 CPU and NVIDIA A6000 GPU, respectively. Xuliang Yu, Yu Qian 0002, Xunzhao Yin, Cheng Zhuo, Liang Zhao 0004 |
DAC | 3 |
| 2025 | FactorHD: A Hyperdimensional Computing Model for Multi-Object Multi-Class Representation and FactorizationabstractNeuro-symbolic artificial intelligence (neurosymbolic AI) excels in logical analysis and reasoning. Hyperdimensional Computing (HDC), a promising braininspired computational model, is integral to neuro-symbolic AI. Various HDC models have been proposed to represent class-instance and class-class relations, but when representing the more complex class-subclass relation, where multiple objects associate different levels of classes and subclasses, they face challenges for factorization, a crucial task for neuro-symbolic AI systems. In this article, we propose FactorHD, a novel HDC model capable of representing and factorizing the complex class-subclass relation efficiently. FactorHD features a symbolic encoding method that embeds an extra memorization clause, preserving more information for multiple objects. In addition, it employs an efficient factorization algorithm that selectively eliminates redundant classes by identifying the memorization clause of the target class. Such model significantly enhances computing efficiency and accuracy in representing and factorizing multiple objects with class-subclass relation, overcoming limitations of existing HDC models such as “superposition catastrophe” and “the problem of 2 “. Evaluations show that FactorHD achieves approximately $5667 \times$ speedup at a representation size of $10^{9}$ compared to existing HDC models. When integrated with the ResNet-18 neural network, FactorHD achieves $92.48 \%$ factorization accuracy on the Cifar-10 dataset. Xuchu Huang, Chenyu Ni, Zheyu Yan, Xunzhao Yin, Cheng Zhuo |
DAC | 6 |
| 2025 | FACAM: Design and Optimization of A Compact Energy Efficient FeFET-Based Analog Content Addressable MemoryabstractContent Addressable Memory (CAM) is known for highly parallel pattern matching capability, which is widely used for data-centric applications and advanced machine learning models that involve associative search tasks. However, most state-of-the-art CAM designs focus on binary/multi-bit CAMs (B/MCAMs) based on CMOS or emerging nonvolatile memories (NVMs), which struggle in scenarios where analog values, rather than discrete levels, need to be stored and searched. Therefore, analog CAMs (ACAMs) offer a promising solution to further increase memory density, improve energy efficiency and extend practical scenarios. Among NVMs, ferroelectric field effect transistors (FeFETs) have emerged as a strong candidate for efficient CAM designs due to the three-terminal structure, high on-off ratio, high OFF resistance and voltage-driven write/read mechanisms. In this paper, we propose FACAM, a compact and energy efficient single-input 2FeFET-1T ACAM cell design, with a two-phase search scheme, that sets the location and width of the matching range through two FeFETs, respectively. We further present a FACAM array which reduces the matchline (ML) voltage swing by shifting ML precharging into the in-cell search operations. We also propose an adaptive scheme to selectively early-terminate second search phase for further search energy optimization. Evaluation results suggest that our proposed FACAM achieves 8.39× and 2.94× energy efficiency compared with the state-of-the-art better 6T-2R ACAM and 2FeFET ACAM. Benchmarking results in deep random forest accelerator show that our approach is 2.14× faster and 7.82× energy efficient than 2FeFET ACAM. Jiahao Cai, Ann Franchesca Laguna, Thomas Kämpfe, Zheyu Yan, Cheng Zhuo, Xunzhao Yin |
ICCAD | 8 |
| 2025 | Accelerating Electro-Thermal Co-Analysis via Coarse-to-Fine Physics-Informed Neural NetworksabstractElectro-thermal coupling has become a concerning issue in 3D integrated circuit (IC) designs. Conventional electro-thermal co-simulation methods rely on iterative solutions of electrical and thermal partial differential equations (PDEs) using numerical techniques, which are computationally expensive and time-consuming. To address this, in this paper, we propose a novel electro-thermal co-analysis framework based on physics-informed neural networks (PINNs) with coarse-to-fine models. The coarse-grained models first predict the electrical potential and temperature distributions of the entire circuit under various boundary conditions, while the fine-grained models provide enhanced resolution for regions of interest. Additionally, we introduce an efficient training strategy that accelerates convergence. Experimental results show that the proposed framework achieves high accuracy with 0.10-0.19% mean relative error and 3-4 orders of magnitude improvements in efficiency compared to the commercial tool. Songyu Sun, Xunzhao Yin, Zhou Jin 0001, Zhiguo Shi 0001, Cheng Zhuo |
ICCAD | 3 |
| 2025 | Invited Paper: Circuit and Architecture Design with Emerging Computing ParadigmsabstractAs emerging computing paradigms push beyond the limitations of traditional CMOS-based computing using Von Neumann architectures, there is a growing need to rethink and extend Electronic Design Automation (EDA) methodologies to support their unique characteristics. These paradigms—including Approximate Computing, In-Memory Computing, Reconfigurable Field-Effect Transistors (RFETs), and Photonic Computing—represent diverse and promising directions beyond conventional digital design. Collectively, they offer transformative potential for achieving significant improvements in energy efficiency, computational speed, and architectural scalability. For example, application-specific approximate computing enables the design of custom arithmetic circuits that exploit application-level error resilience, allowing for optimized accuracy–power–performance–area (PPA) trade-offs in error-tolerant applications. Similarly, processing-in-non-volatile memories, such as those based on Ferroelectric Field-effect Transistors (FeFETs), enhances energy efficiency by enabling analog computation—particularly for operations like matrix multiplication—directly within the memory arrays. The intrinsic polymorphism of RFETs supports compact, multifunctional logic gates and introduces new opportunities for circuit-level obfuscation and security-aware design. Likewise, photonic analog wavefront computing offers substantial gains in latency and energy efficiency by encoding and processing information in the analog optical domain, leveraging phenomena such as diffraction and interference to perform computation at the speed of light. However, they also introduce a host of new challenges in circuit and architecture design, such as vast and irregular design spaces, analog and non-Boolean behavior, and new device-level constraints that existing EDA tools are not capable of handling. To this end, the current article focuses on the development of efficient and robust EDA frameworks that can enable the practical realization of circuits and architectures in these emerging domains. Salim Ullah, Siva Satyendra Sahoo, Can Li 0024, Chao Li 0065, Liu Liu 0023, Tomas Sousa Pereira, Xunzhao Yin, Armin Darjani, Nima Kavand, Chakravarthy Bodla, Rupa Yashaswi Panduga, Aniruddh Holemadlu, Johannes Maly, Jonathan Förste, Samarth Vadia, Xiaobo Sharon Hu, Akash Kumar 0001 |
ICCAD | 8 |
| 2025 | ANAS: Software-hardware co-design of approximate neural network accelerators via neural architecture search
Zheyu Yan, Xunzhao Yin, Lenian He, Cheng Zhuo |
Integr. | 3 |
| 2025 | High-Performance In-Memory Bayesian Inference With Multi-Bit Ferroelectric FETabstractConventional neural network-based machine learning algorithms often encounter difficulties in data-limited scenarios or where interpretability is critical. Conversely, Bayesian inference-based models excel with reliable uncertainty estimates and explainable predictions. Recently, many in-memory computing (IMC) architectures achieve exceptional computing capacity and efficiency for neural network tasks leveraging emerging nonvolatile memory (NVM) technologies. However, their application in Bayesian inference remains limited because the operations in Bayesian inference differ substantially from those in neural networks. In this article, we introduce a compact in-memory Bayesian inference engine with high efficiency and performance utilizing a multi-bit ferroelectric field-effect transistor (FeFET). This design encodes a Bayesian model within a compact FeFETbased crossbar by mapping quantized probabilities to discrete FeFET states. Consequently, the crossbar’s outputs naturally represent the output posteriors of the Bayesian model. Our design facilitates efficient Bayesian inference, accommodating various input types and probability precisions, without additional calculation circuitry. As the first FeFET-based in-memory Bayesian inference engine, our design demonstrates a notable storage density of 26.32 Mb/mm2and a computing efficiency of 581.40 TOPS/W in a representative Bayesian classification task, indicating a 10.7×/43.4× compactness/efficiency improvement compared to the state-of-the-art alternative. Utilizing the proposed Bayesian inference engine, we develop a feature selection system that efficiently addresses a representative NP-hard optimization problem, showcasing our design’s capability and potential to enhance various Bayesian inference-based applications. Test results suggest that our design identifies the essential features, enhancing the model’s performance while reducing its complexity, surpassing the latest implementation in operation speed and algorithm efficiency by 2.9×/2.0×, respectively. Chao Li 0065, Xuchu Huang, Ruibin Mao, Thomas Kämpfe, Kai Ni 0004, Can Li 0024, Xunzhao Yin, Cheng Zhuo |
IEEE Trans. Computers | 10 |
| 2025 | A Scalable 2T-1FeFET-Based Content Addressable Memory Design for Energy Efficient Data SearchabstractContent addressable memory (CAM) is widely used in advanced machine learning models and data-intensive applications for associative search tasks, thanks to the highly parallel pattern matching capability. Most state-of-the-art CAM designs primarily aim to reduce the CAM cell area by utilizing nonvolatile memories (NVMs). However, there has been limited research on optimizing the design and energy efficiency of NVM-based CAMs for practical deployment in edge devices and AI hardware. This article introduces a general compact and energy efficient CAM design scheme that minimizes design overhead by using only one NVM device per cell. Our proposed CAM design realizes both binary CAM (BCAM) and multibit CAM (MCAM) by leveraging the binary and multilevel storage property of NVM devices without additional cell overheads. Additionally, we propose an adaptive matchline (ML) precharge and discharge scheme to further optimize search energy by significantly reducing the ML voltage swing. Ferroelectric field-effect transistors (FeFETs) serve as representative NVMs in our proposed design, and we present a 2T-1FeFET CAM array incorporating a sense amplifier that implements the proposed ML scheme. Evaluation results show that our proposed 2T-1FeFET BCAM design achieves energy efficiency improvements of$6.64\times $/$4.74\times $/$9.14\times $/$3.02\times $compared to CMOS/ReRAM/STT-MRAM/2FeFET BCAM arrays, while 2T-1FeFET MCAM design achieves$8.25\times $/$5.68\times $/$56.35\times $better-energy efficiency compared to ReRAM/3T-1FeFET/1FeFET-1R MACM arrays. Benchmarking results demonstrate that our BCAM/MCAM approach provides$3.2\times $/$3.7\times $and$2.0\times $/$2.2\times $energy-delay product improvement over the 2T-2R and 2FeFET CAM in accelerating query processing applications. Jiahao Cai, Hamza Errahmouni Barkam, Mohsen Imani, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 9 |
| 2025 | CSA-CiM: Enhancing Multifunctional Computing-in-Memory With Configurable Sense AmplifiersabstractComputing-in-memory (CiM) effectively alleviates the memory wall problem faced by traditional von Neumann architectures when handling data-intensive applications. Most CiM arrays employ dedicated sense amplifiers (SAs) to perform specific functions, and prior configurable CiM arrays achieve multifunctionality by stacking multiple SAs with corresponding functions. However, the independent nature of these SAs, particularly the analog-to-digital converter (ADC), results in excessive energy and area consumption. In this article, we propose a configurable multifunctional ferroelectric field effect transistor (FeFET)-based CiM array design, including configurable peripheral circuit with corresponding multifunctionalities and reusable SA components, to reduce energy consumption and latency. The array cells perform logical AND and XNOR operations, and the proposed SA can be configured to operate in either ADC or winner-take-all (WTA) modes, thereby enabling the array to implement both multiplication-accumulation (MAC) and associative search operations. Instead of operating independently, the WTA component within the SA participates as a flash stage in successive approximation register (SAR) conversions in ADC mode, thus enhancing the WTA utilization, energy efficiency and compactness. By integrating the multifunctional CiM array and the configurable SA, our design supports MAC, Hamming-distance computation (HDC), and nearest neighbor search (NNS) operations within the same structure. Compared to existing works, our design achieves energy efficiency improvements of$7.2\times $for MAC,$2.9\times $for HDC, and EDP improvement of$6.4\times $for NNS, respectively. Yuxiao Jiang, Kai Ni 0004, Thomas Kämpfe, Cheng Zhuo, Zheyu Yan, Xunzhao Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2025 | Fast Machine-Learning-Driven Supply Noise-Aware Macromodeling for High-Speed Nonlinear DriversabstractEmerging domains, such as artificial intelligence, 5G mobile, and automotive, are increasingly reliant on high-speed circuits for efficient processing, in which achieving high operating frequencies and data rates is crucial to enable productive data exchange and rapid responses. High-speed data as well as low noise margin in the high-speed serial links call for efficient models of drivers. In this article, we propose a fast machine-learning-driven macromodel for high-speed drivers, which can efficiently capture the nonlinear characteristics of drivers considering dynamic supply noise with low model complexity. A decoupling-superposition strategy is employed to effectively calculate the impact of power supply noise. Additionally, we introduce a piecewise-segmented method for macromodel solving to further enhance the speed of model utilization. Experimental results demonstrate that compared to HSPICE, the proposed macromodel achieves up to$50\times $–$1200\times $speedup while maintaining sufficient accuracy, even for signals with GHz data rate. Songyu Sun, Qi Sun 0002, Xunzhao Yin, Quan Chen 0007, Cheng Zhuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2025 | A Homogeneous FeFET-Based Time-Domain Compute-in-Memory Fabric for Matrix-Vector Multiplication and Associative SearchabstractMatrix-vector multiplication (MVM) and content-based search are two key operations in many machine learning workloads. This article proposes a ferroelectric FET (FeFET) time-domain compute-in-memory (TD-CiM) array that can accelerate both operations in a homogeneous fabric. We demonstrate that 1) the AND and xor/XNOR logic functions required by MVM and content-based search can be realized using a single compute-in-memory (CiM) cell composed of 2FeFETs connected in series; 2) an inverter chain-based TD-CiM array along with a two-phase time-domain computation principle of the TD-CiM can be employed to implement the MVM and content-based search functions; 3) a signal delay-to-digital output conversion can be implemented by associating a loading capacitor with each stage of the inverter chain-based TD-CiM array, ensuring the full digital compatibility; and 4) the proposed 2FeFET cell and inverter chain-based TD-CiM array are robust against FeFET variation according to our comprehensive theoretical and experimental validation. We show how the FeFET TD-CiM can be exploited to accelerate hyperdimensional computing (HDC) and adjusted to process different tasks through dynamic and fine-grained resource allocation. HDC application benchmarking results show that the proposed FeFET-based TD-CiM offers on average$106\times $/$63\times $energy reduction/speedup compared to GPU-based implementation. With more than 8500 TOPS/W energy-efficiency, the proposed FeFET-based TD-CiM exhibits huge potential as a processing fabric for various memory-intensive applications. Xunzhao Yin, Qingrong Huang, Hamza Errahmouni Barkam, Franz Müller 0001, Shan Deng, Alptekin Vardar, Sourav De 0002, Zhouhang Jiang, Mohsen Imani, Ulf Schlichtmann, Xiaobo Sharon Hu, Cheng Zhuo, Thomas Kämpfe, Kai Ni 0004 |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2025 | Ferroelectric Compute-in-Memory Framework for Solving Pure and Mixed Strategy Nash EquilibriumabstractNash equilibrium (NE) is a key concept in game theory, but verifying its existence is NP-complete. Recent advancements proposed quantum NE solvers that identify pure strategy NE solutions (binary solutions) by integrating slack terms into the objective function, known as slack-quadratic unconstrained binary optimization (S-QUBO). However, S-QUBO alters the objective function and can lead to incorrect solutions. Additionally, current solvers only find a limited number of pure strategy NE solutions and cannot address mixed strategy NE (decimal solutions), leaving many solutions unexplored. In this work, we propose C-Nash, a novel ferroelectric compute-in-memory (CiM) framework capable of efficiently addressing both pure and mixed strategy NE solutions. C-Nash consists of 1) a transformation method that transforms quadratic optimization into a MAX-QUBO form without incorporating additional slack variables, thus avoiding objective function changes; 2) A ferroelectric FET (FeFET) based CiM bi-crossbar structure and winner-takes-all (WTA) tree for accelerating the MAX-QUBO form in a single iteration; 3) An efficient operation flow including a rank-based QUBO reformulation algorithm that simplifies the QUBO matrices to reduce hardware overhead, and a two-phase based simulated annealing (SA) logic for finding NE solutions; 4) A FeFET-based crossbar macro for experimental demonstration. Experimental results show that C-Nash increases the success rate for identifying NE solutions by 68.6% while saving$3\times $in chip size. Furthermore, C-Nash can find all pure and mixed NE solutions, unlike D-Wave based quantum approaches which only find some pure strategy NE solutions. Additionally, C-Nash significantly reduces the time-to-solution by up to$157.9\times $/$79.0\times $compared to D-Wave 2000 Q6 and D-Wave Advantage 4.1, respectively. Yu Qian 0002, Ding Huang, Alptekin Vardar, Nellie Laleni, Kai Ni 0004, Thomas Kämpfe, Cheng Zhuo, Xunzhao Yin |
IEEE Trans. Circuits Syst. I Regul. Pap. | 9 |
| 2025 | SenHDC: A 3-D NAND Flash-Based Processing-in-Sensor Hyperdimensional Computing ArchitectureabstractThe rapid growth of the Internet of Things (IoT) promotes vision application deployments in embedded devices. To alleviate the data conversion and transmission overheads in CMOS image sensor (CIS), processing-in-sensor (PIS) has been proposed, enabling computations to occur directly within sensory systems. Meanwhile, brain-inspired hyperdimensional computing (HDC) emerges as a promising computing paradigm well-suited for PIS due to its high accuracy and efficiency in various cognitive tasks. However, HDC requires transforming the sensed signals into long hypervectors (HVs), which imposes significant overhead for PIS systems. Thus, in this article, we propose SenHDC, a hardware-software co-design framework for HDC-based PIS. We introduce a novel hardware-friendly encoding paradigm that eliminates complex HV transformations, coupled with a noise-resilient position HV generation method that mitigates the impact of capacitance load imbalance. We further present an efficient 3-D NAND flash-based compute-in-memory (CIM) hardware design, comprising an encoding module that performs encoding in charge domain and an associative search module that checks the similarity between the input and prestored classes for classification. Experimental results show that SenHDC improves energy efficiency by$40.6\times $and$1.6\times $in the encoding module and associative search module, respectively, compared to the state-of-the-art HDC CIM implementations. Xuchu Huang, Qingrong Huang, Zheyu Yan, Cheng Zhuo, Xunzhao Yin |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | ConvFIFO: A Crossbar Memory PIM Architecture for ConvNets Featuring First-In-First-Out DataflowabstractProcess-in-memory (PIM) architectures based on emerging non-volatile memories (NVMs) have been widely studied for more efficient computation of convolutional neural networks (ConvNets). However, conventional NVM-based PIM suffered from various non-idealities including IR drop, sneak-path currents, analog-to-digital converter (ADC) overhead, device variations and mismatch. In this work, we propose ConvFIFO, a crossbar memory PIM architecture for ConvNets featuring a novel first-in-first-out (FIFO) dataflow. Through the design of FIFO-type input/output buffers, ConvFIFO can maximize the reuse rates of inputs and partial sums to achieve a more balanced trade-off among throughput, accuracy and area/energy consumption. By using SRAM-based FIFO, ConvFIFO further achieves a systolic architecture without the need to move weight data, bypassing the limitation of NVM endurance. Compared to classical NVM-based PIM architectures like ISAAC, ConvFIFO exhibits significant performance improvement in terms of energy consumption $(1.66-3.56 \times)$, latency $(1.69-1.74 \times)$, Ops/W ($4.23-10.17 \times)$ and $\mathrm{Ops}/\mathrm{s} \times \mathrm{mm}^{2} (1.59-1.74 \times)$, benchmarked against a number of common ConvNet models. Liang Zhao 0004, Yu Qian 0002, Fanzi Meng, Xiapeng Xu, Xunzhao Yin, Cheng Zhuo |
ASPDAC | 5 |
| 2024 | FeBiM: Efficient and Compact Bayesian Inference Engine Empowered with Ferroelectric In-Memory ComputingabstractIn scenarios with limited training data or where explainability is crucial, conventional neural network-based machine learning models often face challenges. In contrast, Bayesian inference-based algorithms excel in providing interpretable predictions and reliable uncertainty estimation in these scenarios. While many state-of-the-art in-memory computing (IMC) architectures leverage emerging non-volatile memory (NVM) technologies to offer unparalleled computing capacity and energy efficiency for neural network workloads, their application in Bayesian inference is limited. This is because the core operations in Bayesian inference, i.e., cumulative multiplications of prior and likelihood probabilities, differ significantly from the multiplication-accumulation (MAC) operations common in neural networks, rendering them generally unsuitable for direct implementation in most existing IMC designs. In this paper, we propose FeBiM, an efficient and compact Bayesian inference engine powered by multi-bit ferroelectric field-effect transistor (FeFET)-based IMC. FeBiM effectively encodes the trained probabilities of a Bayesian inference model within a compact FeFET-based crossbar. It maps quantized logarithmic probabilities to discrete FeFET states. As a result, the accumulated outputs of the crossbar naturally represent the posterior probabilities, i.e., the Bayesian inference model's output given a set of observations. This approach enables efficient in-memory Bayesian inference without the need for additional calculation circuitry. As the first FeFET-based in-memory Bayesian inference engine, FeBiM achieves an impressive storage density of 26.32 Mb/mm2 and a computing efficiency of 581.40 TOPS/W in a representative Bayesian classification task. These results demonstrate 10.7×/43.4× improvement in compactness/efficiency compared to the state-of-the-art hardware implementation of Bayesian inference. Chao Li 0065, Ruibin Mao, Can Li 0024, Thomas Kämpfe, Kai Ni 0004, Xunzhao Yin |
DAC | 8 |
| 2024 | C-Nash: A Novel Ferroelectric Computing-in-Memory Architecture for Solving Mixed Strategy Nash EquilibriumabstractThe concept of Nash equilibrium (NE), pivotal within game theory, has garnered widespread attention across numerous industries. However, verifying the existence of NE poses a significant computational challenge, classified as an NP-complete problem. Recent advancements introduced several quantum Nash solvers aimed at identifying pure strategy NE solutions (i.e., binary solutions) by integrating slack terms into the objective function, commonly referred to as slack-quadratic unconstrained binary optimization (S-QUBO). However, incorporation of slack terms into the quadratic optimization results in changes of the objective function, which may cause incorrect solutions. Furthermore, these quantum solvers only identify a limited subset of pure strategy NE solutions, and fail to address mixed strategy NE (i.e., decimal solutions), leaving many solutions undiscovered. In this work, we propose C-Nash, a novel ferroelectric computing-in-memory (CiM) architecture that can efficiently handle both pure and mixed strategy NE solutions. The proposed architecture consists of (i) a transformation method that converts quadratic optimization into a MAX-QUBO form without introducing additional slack variables, thereby avoiding objective function changes; (ii) a ferroelectric FET (FeFET) based bi-crossbar structure for storing payoff matrices and accelerating the core vector-matrix-vector (VMV) multiplications of QUBO form; (iii) A winner-takes-all (WTA) tree implementing the MAX form and a two-phase based simulated annealing (SA) logic for searching NE solutions. Evaluations show that C-Nash has up to 68.6% increase in the success rate for identifying NE solutions, finding all pure and mixed NE solutions rather than only a portion of pure NE solutions, compared to D-Wave based quantum approaches. Moreover, C-Nash boasts a reduction up to 157.9X/79.0X in time-to-solutions compared to D-Wave 2000 Q6 and D-Wave Advantage 4.1, respectively. Yu Qian 0002, Kai Ni 0004, Thomas Kämpfe, Cheng Zhuo, Xunzhao Yin |
DAC | 5 |
| 2024 | HyCiM: A Hybrid Computing-in-Memory QUBO Solver for General Combinatorial Optimization Problems with Inequality ConstraintsabstractComputationally challenging combinatorial optimization problems (COPs) play a fundamental role in various applications. To tackle COPs, many Ising machines and Quadratic Unconstrained Binary Optimization (QUBO) solvers have been proposed, which typically involve direct transformation of COPs into Ising models or equivalent QUBO forms (D-QUBO). However, when addressing COPs with inequality constraints, this D-QUBO approach introduces numerous extra auxiliary variables, resulting in a substantially larger search space, increased hardware costs, and reduced solving efficiency. In this work, we propose HyCiM, a novel hybrid computing-inmemory (CiM) based QUBO solver framework, designed to overcome aforementioned challenges. The proposed framework consists of (i) an innovative transformation method (first to our known) that converts COPs with inequality constraints into an inequality-QUBO form, thus eliminating the need of expensive auxiliary variables and associated calculations; (ii) "inequality filter", a ferroelectric FET (FeFET)-based CiM circuit that accelerates the inequality evaluation, and filters out infeasible input configurations; (iii) a FeFET-based CiM annealer that is capable of approaching global solutions of COPs via iterative QUBO computations within a simulated annealing process. The evaluation results show that HyCiM drastically narrows down the search space, eliminating 2100 to 22536 infeasible input configurations compared to the conventional D-QUBO approach. Consequently, the narrowed search space, reduced to 2100 feasible input configurations, leads to a substantial hardware area overhead reduction, ranging from 88.06% to 99.96%. Additionally, HyCiM consistently exhibits a high solving efficiency, achieving a remarkable average success rate of 98.54%, whereas D-QUBO implementatoin shows only 10.75%. Yu Qian 0002, Kai Ni 0004, Alptekin Vardar, Thomas Kämpfe, Xunzhao Yin |
DAC | 6 |
| 2024 | Energy Efficient Dual Designs of FeFET-Based Analog In-Memory Computing with Inherent Shift-Add CapabilityabstractIn-memory computing (IMC) architecture emerges as a promising paradigm, improving the energy efficiency of multiply-and-accumulate (MAC) operations within deep neural networks (DNNs) by integrating the parallel computations within the memory arrays. Various high-precision analog IMC array designs have been developed based on both SRAM and emerging non-volatile memories (NVMs). These designs perform MAC operations of partial input and weight, with the corresponding partial products then fed into shift-add circuitry to produce the final MAC results. However, existing works often present intricate shift-add process for weight. The traditional digital shift-add process is limited in throughput due to time-multiplexing of ADCs, and advancing the shift-add process to the analog domain necessitates customized circuit implementations, resulting in compromises in energy and area efficiency. Furthermore, the joint optimization of the partial MAC operations and the weight shift-add process is rarely explored. In this paper, we propose novel, energy efficient dual designs of ferroelectric FET (FeFET) based high precision analog IMC featuring inherent shift-add capability. We introduce a FeFET based IMC paradigm that performs partial MAC in each column, and inherently integrates the shift-add process for 4-bit weights by leveraging FeFET's analog storage characteristics. This paradigm supports both 2's complement mode (2CM) and non-2's complement mode (N2CM) MAC, thereby offering flexible support for 4-/8-bit weight data in 2's complement format. Building upon this paradigm, we propose novel FeFET based dual designs, CurFe for the current mode and ChgFe for the charge mode, to accommodate the high precision analog domain IMC architecture. Evaluation results at circuit and system levels indicate that the circuit/system-level energy efficiency of the proposed FeFET-based analog IMC is 1.56×/1.37× higher when compared to the state-of-the-art analog IMC designs. Qingrong Huang, Yu Qian 0002, Kai Ni 0004, Thomas Kämpfe, Xunzhao Yin |
DAC | 6 |
| 2024 | Computational and Storage Efficient Quadratic Neurons for Deep Neural NetworksabstractDeep neural networks (DNNs) have been widely deployed across diverse domains such as computer vision and natural language processing. However, the impressive accomplishments of DNNs have been realized alongside extensive computational demands, thereby impeding their applicability on resource-constrained devices. To address this challenge, many researchers have been focusing on basic neuron structures, the fundamental building blocks of neural networks, to alleviate the computational and storage cost. In this work, an efficient quadratic neuron architecture distinguished by its enhanced utilization of second-order computational information is introduced. By virtue of their better expressivity, DNNs employing the proposed quadratic neurons can attain similar accuracy with fewer neurons and computational cost. Experimental results have demonstrated that the proposed quadratic neuron structure exhibits superior computational and storage efficiency across various tasks when compared with both linear and non-linear neurons in prior work. Chuangtao Chen 0001, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Ulf Schlichtmann, Bing Li 0005 |
DATE | 3 |
| 2024 | A FeFET-based Time-Domain Associative Memory for Multi-bit Similarity ComputationabstractThe exponential growth of data across various domains of human society necessitates the rapid and efficient data processing. In many contemporary data-intensive applications, similarity computation (SC) is one of the most fundamental and indispensable operations. In recent years, In-memory computing (IMC) architectures have been designed to accelerate SC by reducing data movement costs, however, they encounter challenges with signal domain conversion, variation sensitivity, and limited precision. This paper proposes a ferroelectric FET (FeFET) based time-domain (TD) associative memory (AM) for energy efficient SC. Such TD design can convert its output (i.e., time interval) to digits with relatively simple sensing circuitry thus saves large amount of area and energy compared with conventional IMC designs that process analog voltage/current signals. The variable-capacitance (VC) delay chain structure in our design supports quantitative SC and enhances robustness against variations. Furthermore, by exploiting multi-domain ferroelctric FET (FeFET), our design is capable of performing SC on vectors with multi-bit element, enabling support for higher-precision algorithms. Simulation results show that the proposed TD-AM achieves 13.8x/1.47x energy saving of our design compared to CMOS/NVM based TD-IMC designs. Additionally, our design exhibits good robustness in monte carlo simulation with variation extracted from experimental measurements. Investigation on precision of hyperdimensional computing (HDC) show that higher element precision reduces the size of HDC model when considering to achieve same accuracy, indicating an improved efficiency. Benchmarkings against GPU demonstrate in general 2/3 orders of magnitude speedup/energy efficiency improvement of our design. Our proposed multi-bit TD-AM promises energy-efficient quantitative SC for diverse intensive data processing application, especially in energy-constrained scenarios. Qingrong Huang, Hamza Errahmouni Barkam, Jianyi Yang 0003, Thomas Kämpfe, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Mohsen Imani, Cheng Zhuo, Xunzhao Yin |
DATE | 12 |
| 2024 | Class-Aware Pruning for Efficient Neural NetworksabstractDeep neural networks (DNNs) have demonstrated remarkable success in various fields. However, the large number of floating-point operations (FLOPs) in DNNs poses challenges for their deployment in resource-constrained applications, e.g., edge devices. To address the problem, pruning has been introduced to reduce the computational cost in executing DNNs. Previous pruning strategies are based on weight values, gradient values and activation outputs. Different from previous pruning solutions, in this paper, we propose a class-aware pruning technique to compress DNNs, which provides a novel perspective to reduce the computational cost of DNNs. In each iteration, the neural network training is modified to facilitate the class-aware pruning. Afterwards, the importance of filters with respect to the number of classes is evaluated. The filters that are only important for a few number of classes are removed. The neural network is then retrained to compensate for the incurred accuracy loss. The pruning iterations end until no filter can be removed anymore, indicating that the remaining filters are very important for many classes. This pruning technique outperforms previous pruning solutions in terms of accuracy, pruning ratio and the reduction of FLOPs. Experimental results confirm that this class-aware pruning technique can significantly reduce the number of weights and FLOPs, while maintaining a high inference accuracy. Our code is available at https://github.com/HWAI-TUDa/Class-Aware-Pruning Mengnan Jiang, Jingcun Wang, Amro Eldebiky, Xunzhao Yin, Cheng Zhuo, Ing-Chao Lin, Grace Li Zhang |
DATE | 4 |
| 2024 | OplixNet: Towards Area-Efficient Optical Split-Complex Networks with Real-to-Complex Data Assignment and Knowledge DistillationabstractHaving the potential for high speed, high throughput, and low energy cost, optical neural networks (ONN s) have emerged as a promising candidate for accelerating deep learning tasks. In conventional ONNs, light amplitudes are modulated at the input and detected at the output. However, the light phases are still ignored in conventional structures, although they can also carry information for computing. To address this issue, in this paper, we propose a framework called OplixNet to compress the areas of ONNs by modulating input image data into the amplitudes and phase parts of light signals. The input and output parts of the ONN s are redesigned to make full use of both amplitude and phase information. Moreover, mutual learning across different ONN structures is introduced to maintain the accuracy. Experimental results demonstrate that the proposed framework significantly reduces the areas of ONNs with the accuracy within an acceptable range. For instance, 75.03 % area is reduced with a 0.33% accuracy decrease on fully connected neural network (FCNN) and 74.88% area is reduced with a 2.38% accuracy decrease on ResNet-32. Ruidi Qiu, Amro Eldebiky, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Ulf Schlichtmann, Bing Li 0005 |
DATE | 4 |
| 2024 | FeReX: A Reconfigurable Design of Multi-Bit Ferroelectric Compute-in-Memory for Nearest Neighbor SearchabstractRapid advancements in artificial intelligence have given rise to transformative models, profoundly impacting our lives. These models demand massive volumes of data to operate effectively, exacerbating the data-transfer bottleneck inherent in the conventional von-Neumann architecture. Compute-in-memory (CIM), a novel computing paradigm, tackles these issues by seam-lessly embedding in-memory search functions, thereby obviating the need for data transfers. However, existing non-volatile memory (NVM)-based accelerators are application specific. During the similarity based associative search operation, they only support a single, specific distance metric, such as Hamming, Manhattan, or Euclidean distance in measuring the query against the stored data, calling for reconfigurable in-memory solutions adaptable to various applications. To overcome such a limitation, in this paper, we present FeReX, a reconfigurable associative memory (AM) that accommodates various distance metrics including Hamming, Manhattan, and Euclidean distances. Leveraging multi-bit ferroelectric field-effect transistors (FeFETs) as the proxy and a hardware-software co-design approach, we introduce a constrained satisfaction problem (CSP)-based method to automate AM search input voltage and stored voltage configurations for different distance based search functions. Device-circuit co-simulations first validate the effectiveness of the proposed FeReX methodology for reconfigurable search distance functions. Then, we benchmark FeReX in the context of k-nearest neighbor (KNN) and hyperdimensional computing (HDC), which highlights the robustness of FeReX and demonstrates up to 250× speedup and 104energy savings compared with GPU. Che-Kai Liu, Chao Li 0065, Ruibin Mao, Jianyi Yang 0003, Thomas Kämpfe, Mohsen Imani, Can Li 0024, Cheng Zhuo, Xunzhao Yin |
DATE | 10 |
| 2024 | Reconfigurable Frequency Multipliers Based on Complementary Ferroelectric TransistorsabstractFrequency multipliers, a class of essential electronic components, play a pivotal role in contemporary signal processing and communication systems. They serve as crucial building blocks for generating high-frequency signals by multiplying the frequency of an input signal. However, traditional frequency multipliers that rely on nonlinear devices often require energy- and area-consuming filtering and amplification circuits, and emerging designs based on an ambipolar ferroelectric transistor require costly non-trivial characteristic tuning or complex technology process. In this paper, we show that a pair of standard ferroelectric field effect transistors (FeFETs) can be used to build compact frequency multipliers without aforementioned technology issues. By leveraging the tunable parabolic shape of the 2FeFET structures' transfer characteristics, we propose four reconfigurable frequency multipliers, which can switch between signal transmission and frequency doubling. Furthermore, based on the 2FeFET structures, we propose four frequency multipliers that realize triple, quadruple frequency modes, elucidating a scalable methodology to generate more multiplication harmonics of the input frequency. Performance metrics such as maximum operating frequency, power, etc., are evaluated and compared with existing works. We also implement a practical case of frequency modulation scheme based on the proposed reconfigurable multipliers without additional devices. Our work provides a novel path of scalable and reconfigurable frequency multiplier designs based on devices that have characteristics similar to FeFETs, and show that FeFETs are a promising candidate for signal processing and communication systems in terms of maximum operating frequency and power. Jianyi Yang 0003, Cheng Zhuo, Thomas Kämpfe, Kai Ni 0004, Xunzhao Yin |
DATE | 6 |
| 2024 | Low Power and Temperature- Resilient Compute-In-Memory Based on Subthreshold-FeFETabstractCompute-in-memory (CiM) is a promising solution for addressing the challenges of artificial intelligence (AI) and the Internet of Things (IoT) hardware such as “memory wall” issue. Specifically, CiM employing nonvolatile memory (NVM) devices in a crossbar structure can efficiently accelerate multiply-accumulation (MAC) computation, a crucial operator in neural networks among various AI models. Low power CiM designs are thus highly desired for further energy efficiency optimization on AI models. Ferroelectric FET (FeFET), an emerging device, is attractive for building ultra-low power CiM array due to CMOS compatibility, high ION /$I$O F F ratio, etc. Recent studies have explored FeFET based CiM designs that achieve low power consumption. Nevertheless, subthreshold-operated FeFETs, where the operating voltages are scaled down to the subthreshold region to reduce array power consumption, are particularly vulnerable to temperature drift, leading to accuracy degradation. To address this challenge, we propose a temperature-resilient 2T-1FeFET CiM design that performs MAC operations reliably at subthreahold region from 0°C to 85°C, while consuming ultra-low power. Benchmarked against the VGG neural network architecture running the CIFAR-10 dataset, the proposed 2T1FeFET CiM design achieves 89.45% CIFAR-10 test accuracy. Compared to previous FeFET based CiM designs, it exhibits immunity to temperature drift at an 8-bit wordlength scale, and achieves better energy efficiency with 2866 TOPS/W. Xuchu Huang, Jianyi Yang 0003, Kai Ni 0004, Hussam Amrouch, Cheng Zhuo, Xunzhao Yin |
DATE | 7 |
| 2024 | BasisN: Reprogramming-Free RRAM-Based In-Memory-Computing by Basis Combination for Deep Neural NetworksabstractDeep neural networks (DNNs) have made breakthroughs in various fields including image recognition and language processing. DNNs execute hundreds of millions of multiply-and-accumulate (MAC) operations. To efficiently accelerate such computations, analog in-memory-computing platforms have emerged leveraging emerging devices such as resistive RAM (RRAM). However, such accelerators face the hurdle of being required to have sufficient on-chip crossbars to hold all the weights of a DNN. Otherwise, RRAM cells in the crossbars need to be reprogramed to process further layers, which causes huge time/energy overhead due to the extremely slow writing and verification of the RRAM cells. As a result, it is still not possible to deploy such accelerators to process large-scale DNNs in industry. To address this problem, we propose the BasisN framework to accelerate DNNs on any number of available crossbars without reprogramming. BasisN introduces a novel representation of the kernels in DNN layers as combinations of global basis vectors shared between all layers with quantized coefficients. These basis vectors are written to crossbars only once and used for the computations of all layers with marginal hardware modification. BasisN also provides a novel training approach to enhance computation parallelization with the global basis vectors and optimize the coefficients to construct the kernels. Experimental results demonstrate that cycles per inference and energy-delay product were reduced to below 1% compared with applying reprogramming on crossbars in processing large-scale DNNs such as DenseNet and ResNet on ImageNet and CIFAR100 datasets, while the training and hardware costs are negligible. Amro Eldebiky, Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Ing-Chao Lin, Ulf Schlichtmann, Bing Li 0005 |
ICCAD | 3 |
| 2024 | TAP-CAM: A Tunable Approximate Matching Engine based on Ferroelectric Content Addressable MemoryabstractPattern search is crucial in numerous analytic applications for retrieving data entries akin to the query. Content Addressable Memories (CAMs), an in-memory computing fabric, directly compare input queries with stored entries through embedded comparison logic, facilitating fast parallel pattern search in memory. While conventional CAM designs offer exact match functionality, they are inadequate for meeting the approximate search needs of emerging data-intensive applications. Some recent CAM designs propose approximate matching functions, but they face limitations such as excessively large cell area or the inability to precisely control the degree of approximation. In this paper, we propose TAP-CAM, a novel ferroelectric field effect transistor (FeFET) based ternary CAM (TCAM) capable of both exact and tunable approximate matching. TAP-CAM employs a compact 2FeFET-2R cell structure as the entry storage unit, and similarities in Hamming distances between input queries and stored entries are measured using an evaluation transistor associated with the matchline of CAM array. The operation, robustness and performance of the proposed design at array level have been discussed and evaluated, respectively. We conduct a case study of K-nearest neighbor (KNN) search to benchmark the proposed TAP-CAM at application level. Results demonstrate that compared to 16T CMOS CAM with exact match functionality, TAP-CAM achieves a 16.95× energy improvement, along with a 3.06% accuracy enhancement. Compared to 2FeFET TCAM with approximate match functionality, TAP-CAM achieves a 6.78× energy improvement. Chenyu Ni, Che-Kai Liu, Liu Liu 0023, Mohsen Imani, Thomas Kämpfe, Kai Ni 0004, Michael T. Niemier, Xiaobo Sharon Hu, Cheng Zhuo, Xunzhao Yin |
ICCAD | 11 |
| 2024 | TReCiM: Lower Power and Temperature-Resilient Multibit 2FeFET-1T Compute-in-Memory DesignabstractCompute-in-memory (CiM) emerges as a promising solution to solve hardware challenges in artificial intelligence (AI) and the Internet of Things (IoT), particularly addressing the "memory wall" issue. By utilizing nonvolatile memory (NVM) devices in a crossbar structure, CiM efficiently accelerates multiplyaccumulate (MAC) computations, the crucial operations in neural networks and other AI models. Among various NVM devices, Ferroelectric FET (FeFET) is particularly appealing for ultra-low-power CiM arrays due to its CMOS compatibility, voltage-driven write/read mechanisms and high ION/IOFF ratio. Moreover, subthreshold-operated FeFETs, which operate at scaling voltages in the subthreshold region, can further minimize the power consumption of CiM array. However, subthreshold-FeFETs are susceptible to temperature drift, resulting in computation accuracy degradation. Existing solutions exhibit weak temperature resilience at larger array size and only support 1-bit. In this paper, we propose TReCiM, an ultra-low-power temperature-resilient multibit 2FeFET-1T CiM design that reliably performs MAC operations in the subthreshold-FeFET region with temperature ranging from 0°C to 85°C at scale. We benchmark our design using NeuroSim framework in the context of VGG-8 neural network architecture running the CIFAR-10 dataset. Benchmarking results suggest that when considering temperature drift impact, our proposed TReCiM array achieves 91.31% accuracy, with 1.86% accuracy improvement compared to existing 1-bit 2T-1FeFET CiM array. Furthermore, our proposed design achieves 48.03 TOPS/W energy efficiency at system level, comparable to existing designs with smaller technology feature sizes. Thomas Kämpfe, Kai Ni 0004, Hussam Amrouch, Cheng Zhuo, Xunzhao Yin |
ICCAD | 6 |
| 2024 | A Survey on Approximate Multiplier Designs for Energy Efficiency: From Algorithms to CircuitsabstractGiven the stringent requirements of energy efficiency for Internet-of-Things edge devices, approximate multipliers, as a basic component of many processors and accelerators, have been constantly proposed and studied for decades, especially in error-resilient applications. The computation error and energy efficiency largely depend on how and where the approximation is introduced into a design. Thus, this article aims to provide a comprehensive review of the approximation techniques in multiplier designs ranging from algorithms and architectures to circuits. We have implemented representative approximate multiplier designs in each category to understand the impact of the design techniques on accuracy and efficiency. The designs can then be effectively deployed in high-level applications, such as machine learning, to gain energy efficiency at the cost of slight accuracy loss. Chuangtao Chen 0001, Weihua Xiao, Xuan Wang 0027, Chenyi Wen, Jie Han 0001, Xunzhao Yin, Weikang Qian, Cheng Zhuo |
ACM Trans. Design Autom. Electr. Syst. | 7 |
| 2024 | Multibit Content Addressable Memory Design and Optimization Based on 3-D nand-Compatible IGZO FlashabstractContent addressable memory (CAM) has been employed in various data-intensive tasks for its parallel pattern-matching capability. To enhance the density and efficiency of CAMs, emerging nonvolatile memory (NVM) technologies have been exploited in the CAM designs. Recently, the multilevel cell (MLC) characteristics of NVMs have been utilized in several analog and multibit CAM designs, achieving higher density than conventional binary/ternary CAM designs. However, these analog and multibit CAM designs are built with the practical experience of circuit designers, lacking a general analog/multibit design methodology. In this article, we propose a general and effective design and optimization scheme for multibit CAM, using a novel 3-D nand-compatible amorphous indium–gallium–zinc–oxide (IGZO) flash as a proxy of three-terminal NVM devices. The proposed scheme encodes the multibit data into the flash devices, enabling the 3-D nand flash array to operate as an ultradense nand or nor CAM without significant structural change. For further performance optimization, we propose a design space exploration scheme for optimal CAM parameters. Evaluation results suggest that the CAM design based on our proposed design and optimization scheme achieves over 35$\times$area per bit saving compared with the representative ferroelectric field effect transistor (FeFET)-based multibit CAM, and a 38.1$\times$energy-delay-area product (EDAP) improvement over the state-of-the-art analog CAM, respectively. Chao Li 0065, Chen Sun 0010, Jianyi Yang 0003, Kai Ni 0004, Cheng Zhuo, Xunzhao Yin |
IEEE Trans. Very Large Scale Integr. Syst. | 7 |
| 2024 | Enhancing ConvNets With ConvFIFO: A Crossbar PIM Architecture Based on Kernel-Stationary First-In-First-Out DataflowabstractConvolutional neural networks (ConvNets) have long been the model of choice for computer vision (CV) problems and gained renewed traction lately. In order to compute ConvNets more efficiently, process-in-memory (PIM) architectures based on emerging non-volatile memories (NVMs) such as RRAM have been widely studied. However, conventional NVM-based PIM suffered from various non-idealities including IR drop, sneak-path currents, large analog-to-digital converter (ADC) overhead, device variations, circuits mismatch, and error propagation. In this work, we propose ConvFIFO, a crossbar-memory-based PIM architecture for ConvNets featuring a kernel-stationary dataflow. Through the design of FIFO-type input and output buffers, smaller row-activation parallelism, and more compact ADCs, ConvFIFO can maximize the reuse rates of inputs and partial sums to achieve a more balanced trade-off among throughput, accuracy, and area/energy consumption. Using SRAM-based FIFO as the input/output buffer, ConvFIFO achieves a systolic architecture without the need to move weight data, bypassing the limitation of NVM endurance and minimizing the movement of partial sums. Moreover, the FIFO nature of the dataflow allows flexible pipeline design and load balancing. Compared to classical NVM-based PIM architectures such as ISAAC, ConvFIFO exhibits significant performance enhancement for various ConvNet models, showing 1.66–$1.69\times $/1.69–$1.74\times $/4.23–$4.79\times $/1.59–$1.74\times $improvement in terms of energy consumption, latency, Ops/W, and Ops/s$\times $mm2, respectively. Compared to GPUs, ConvFIFO exhibits only an average accuracy loss of 1.82% during inference. Yu Qian 0002, Liang Zhao 0004, Fanzi Meng, Xiapeng Xu, Cheng Zhuo, Xunzhao Yin |
IEEE Trans. Very Large Scale Integr. Syst. | 6 |
| 2023 | Approximate Floating-Point FFT Design with Wide Precision-Range and High Energy EfficiencyabstractFast Fourier Transform (FFT) is a key digital signal processing algorithm that is widely deployed in mobile and portable devices. Recently, with the popularity of human perception related tasks, it is noted that the requirements of full precision and exactness are not always necessary for FFT computation. We propose a top-down approximate Floating-Point FFT design methodology to fully exploit the error-tolerance nature of the FFT algorithm. An efficient error modeling of the configurable approximate multiplier is proposed to link the multiplier approximation to the FFT algorithm precision. Then an approximation optimization flow is formulated to maximize the energy efficiency. Experimental results show that the proposed approximate FFT can achieve up to 52% Area-Delay-Product improvement and 23% energy saving when compared to the exact FFT. The proposed approximate FFT is also found to cover almost 2X wider precision range with higher energy efficiency in comparison with the prior state-of-the-art approximate FFT. Chenyi Wen, Xunzhao Yin, Cheng Zhuo |
ASP-DAC | 3 |
| 2023 | SteppingNet: A Stepping Neural Network with Incremental Accuracy EnhancementabstractDeep neural networks (DNNs) have successfully been applied in many fields in the past decades. However, the in-creasing number of multiply-and-accumulate (MAC) operations in DNNs prevents their application in resource-constrained and resource-varying platforms, e.g., mobile phones and autonomous vehicles. In such platforms, neural networks need to provide ac-ceptable results quickly and the accuracy of the results should be able to be enhanced dynamically according to the computational resources available in the computing system. To address these challenges, we propose a design framework called SteppingNet. SteppingNet constructs a series of sub nets whose accuracy is incrementally enhanced as more MAC operations become avail-able. Therefore, this design allows a trade-off between accuracy and latency. In addition, the larger sub nets in SteppingNet are built upon smaller subnets, so that the results of the latter can directly be reused in the former without recomputation. This property allows SteppingNet to decide on-the-fly whether to enhance the inference accuracy by executing further MAC operations. Experimental results demonstrate that SteppingNet provides an effective incremental accuracy improvement and its inference accuracy consistently outperforms the state-of-the-art work under the same limit of computational resources. Grace Li Zhang, Xunzhao Yin, Cheng Zhuo, Huaxi Gu, Bing Li 0005, Ulf Schlichtmann |
DATE | 3 |
| 2023 | SEE-MCAM: Scalable Multi-Bit FeFET Content Addressable Memories for Energy Efficient Associative SearchabstractArtificial intelligence has made remarkable advancements in recent years, leading to the development of algorithms and models capable of handling ever-increasing amounts of data. The computational demands of these algorithms necessitate circuit and architecture designs that go beyond the von-Neumann paradigm. Content addressable memories (CAMs), which implement parallel associative search functionality within memory blocks to overcome the memory wall bottleneck, have proven to be effective for data-intensive tasks. While current CAM designs have achieved higher storage density and energy efficiency than their CMOS-based counterparts by leveraging emerging non-volatile memories (NVM), most of these implementations are limited to binary storage cells. In this work, we propose SEE-MCAM, scalable and compact multi-bit CAM (MCAM) designs that utilize the three-terminal ferroelectric FET (FeFET) as the proxy. By exploiting the multi-level-cell characteristics of FeFETs, our proposed SEE-MCAM designs enable multi-bit associative search functions and achieve better energy efficiency and performance than existing FeFET-based CAM designs. We validated the functionality of our proposed designs by achieving 3 bits per cell CAM functionality, resulting in 3x improvement in storage density. The area per bit of the proposed SEE-MCAM cell is 8% of the conventional CMOS CAM. We thoroughly investigated the scalability and robustness of the proposed design. Evaluation results suggest that the proposed 2FeFET-1 T SEE-MCAM achieves 9.8× more energy efficiency and 1.6× less search latency compared to the CMOS CAM, respectively. When compared to existing MCAM designs, the proposed SEE-MCAM can achieve 8.7× and 4.9× more energy efficiency than ReRAM-based and FeFET-based MCAMs, respectively. Benchmarking results show that our approach provides up to 3 orders of magnitude improvement in speedup and energy efficiency over a GPU implementation in accelerating a novel quantized hyperdimensional computing (HDC) application. Shengxi Shou, Che-Kai Liu, Sanggeon Yun, Zishen Wan, Kai Ni 0004, Mohsen Imani, Xiaobo Sharon Hu, Jianyi Yang 0003, Cheng Zhuo, Xunzhao Yin |
ICCAD | 10 |
| 2023 | Breaking the energy-efficiency barriers for smart sensing applications with "Sensing with Computing" architectures
Xinghua Yang, Zheyu Liu, Kechao Tang, Xunzhao Yin, Cheng Zhuo, Qi Wei 0001, Fei Qiao |
Sci. China Inf. Sci. | 4 |
| 2023 | LMM: A Fixed-Point Linear Mapping Based Approximate Multiplier for IoT
Chenyi Wen, Xunzhao Yin, Cheng Zhuo |
J. Comput. Sci. Technol. | 3 |
| 2023 | BRoCoM: A Bayesian Framework for Robust Computing on Memristor CrossbarabstractMemristor crossbar arrays are considered to be a promising platform for neuromorphic computing. To deploy a trained neural network (NN) model on memristor crossbars, memristors need to be programmed to the corresponding weight values. In fact, due to device-based process variation and noise, deviations of the stored weights from the trained weights are inevitable, thereby causing the degradation of the actual inference performance. This article proposes a unified Bayesian inference-based framework, BRoCoM, which connects device nonidealities and algorithmic training together for robust computing on memristor crossbars. BRoCoM is able to incorporate different levels of nonidealities into prior weight distribution, and transform robustness optimization to Bayesian NN (BNN) training, the weights of NNs are optimized to accommodate uncertainties and minimize inference degradation. Experimental results confirm the capability of the proposed BRoCoM to achieve stable inference performance while tolerating the nonideal effects of process variation and noise. Qingrong Huang, Grace Li Zhang, Xunzhao Yin, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2023 | FeFET-Based In-Memory Hyperdimensional Encoding DesignabstractThe data explosion of Internet of Things (IoT) and machine learning tasks raises a great demand on highly efficient computing hardware and paradigms. Brain-inspired hyperdimensional computing (HDC) is becoming a promising computing paradigm, which encodes data as hypervectors with homogeneous elements instead of numbers, and can perform learning/classification tasks through simple logical or arithmetic operations on the encoded hypervectors. Therefore HDC has much lower computational complexity than conventional computational models such as neural networks. However, due to its high-dimensional data representation, processing, and encoding hypervectors in conventional Von–Neumann architectures (e.g., CPU and GPU) requires a large amount of energy- and time-consuming data transfer, thus weakening its efficiency benefiting from low complexity. In this article, we proposed an ultralow power and fast computing-in-memory (CiM) design based on nonvolatile (NV) ferroelectric FET (FeFET) for HDC encoding. The proposed design mainly support hyperdimensional bit-wise XOR and parallel majority vote (MAJ) operations for HDC encoding, which are implemented by FeFET-based memories together with CMOS peripheral circuits. The 1FeFET1T-based memory cell effectively mitigates the impact of transistor variations on the operation. A highly parallel and pipelined computing workflow of the proposed design further boosts the energy efficiency and performance with a negligible extra area overhead. Experimental results demonstrate that our proposed design achieves$5.04\times $energy efficiency improvement over other CiM designs for HDC encoding. Qingrong Huang, Kai Ni 0004, Mohsen Imani, Cheng Zhuo, Xunzhao Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 6 |
| 2023 | Ferroelectric Ternary Content Addressable Memories for Energy-Efficient Associative SearchabstractA fast and efficient search function across the database has been a core component for a number of data-intensive tasks in machine learning, IoT applications, and inference. However, the conventional digital machines implementing the search functionality with repetitive arithmetic operations suffer from the energy efficiency and performance degradation due to the significant data transfer between the storage and processing units in the Von Neumann architecture. Ternary content addressable memories (TCAMs) are an essential hardware form of computing-in-memory (CiM) designs that aim to overcome the data transfer bottlenecks by implementing the parallel associative search function within the memory blocks. While most state-of-the-art TCAM designs focus on improving the information density by harnessing compact nonvolatile memories (NVMs), little efforts have been spent on optimizing the energy efficiency of the NVM-based TCAM. In this article, by exploiting the ferroelectric FET (FeFET) as a representative NVM, we propose an NOR-type 2FeFET-1T and an NAND-type 2FeFET-2T TCAM designs that enable highly energy-efficient associative search by reducing the associated precharge overheads. We then propose a hybrid ferroelectric NAND-NOR (HFNN) TCAM design to further improve the energy efficiency. An HFNN-based segmented architecture is proposed to reduce the search delay and energy by search operation pipeline. Evaluation results suggest that the proposed 2FeFET-1T, 2FeFET-2T and HFNN TCAM design consume$3.03\times $,$8.08\times $, and$226.92\times $less search energy than the conventional 16T complementary metal oxide semiconductor (CMOS) TCAM, respectively. Application benchmarking shows that our proposed 2FeFET-1T/2FeFET-2T/HFNN TCAM can save, on average, 45.2%/50.6%/57.5% the GPU energy consumption as compared to the conventional GPU. Xunzhao Yin, Yu Qian 0002, Mohsen Imani, Kai Ni 0004, Chao Li 0065, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 1 |
| 2023 | Design of Ultracompact Content Addressable Memory Exploiting 1T-1MTJ CellabstractContent addressable memories (CAMs) are a promising category of computing-in-memory (CiM) elements that can perform highly parallel and efficient search operations for routers, pattern matching, and other data-intensive applications. Various magnetic tunnel junction (MTJ)-based CAM designs have been proposed to realize zero standby power and high-performance search. However, due to the relatively small tunnel magneto-resistance (TMR) ratio, MTJ-based CAMs require extra transistors and differential MTJ branches to distinguish between the parallel and anti-parallel resistance states, resulting in significant area and energy overhead. In this article, we propose a device-circuit co-design approach for an ultracompact CAM design by only exploiting a 1T-1MTJ structure in each cell. We propose a 2-step search scheme to enable the parallel in-memory search operation across the proposed CAM array and demonstrate the sufficient sensing margin of the array in a successful search operation. Evaluation results suggest that our proposed 1T-1MTJ-based CAM design improves$179\times /301\times $area efficiency compared with the state-of-the-art 15T-4MTJ/20T-6MTJ CAM design. Application benchmarking on hyperdimensional computing (HDC) inference shows a$54.6\times /12.8\times $speedup compared with GPU/20T-6MTJ CAM-based approaches. Cheng Zhuo, Kai Ni 0004, Mohsen Imani, Yuxuan Luo 0001, Shaodi Wang, Deming Zhang, Xunzhao Yin |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2023 | A Reconfigurable FeFET Content Addressable Memory for Multi-State Hamming DistanceabstractPattern searches, a key operation in many data analytic applications, often deal with data represented by multiple states per dimension. However, hash tables, a common software-based pattern search approach, require a large amount of additional memory, and thus, are limited by the memory wall. A hardware-based solution is to use content-addressable memories (CAMs) that support fast associative searches in parallel. Ternary CAMs (TCAMs) support bit-wise Hamming distance (HD) based searches. Detecting the HD of vectors with multiple states per dimension (i.e., multi-state Hamming distance (MSHD)) can be implemented on TCAMs with one-hot encoding, but requires one TCAM cell per state, leading to a higher area, latency, and energy overhead. We propose a Ferroelectric FET (FeFET)-based multi-state CAM design, MHCAM, which implements MSHD searches in a dense FeFET-based memory array. MHCAM only uses$\lceil log_{2} s \rceil ~2$FeFET CAM cells to represent$s$states or symbols per dimension, and can be reconfigured to 2-bit/4-bit/6-bit/8-bit dimensions. A low-cost sensing circuit with matchline voltage scaling technique is introduced to perform both exact match and threshold match. We use DNA and protein pre-alignment filtering as application case studies to evaluate the application-level benefit of MHCAM. DNA and protein pre-alignment filtering achieve$3.8\times /4.7\times $speedup and$1.7\times /1.8\times $energy improvement compared with the state-of-the-art 2FeFET TCAM-based implementation. Liu Liu 0023, Ann Franchesca Laguna, Ramin Rajaei, Mohammad Mehdi Sharifi, Arman Kazemi, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 6 |
| 2023 | Worst-case Power Integrity Prediction Using Convolutional Neural NetworkabstractPower integrity analysis is an essential step in power distribution network (PDN) sign-off to ensure the performance and reliability of chips. However, with the growing PDN size and increasing scenarios to be validated, it becomes very time- and resource-consuming to conduct full-stack PDN simulation to check the power integrity for different test vectors. Recently, various works have proposed machine learning–based methods for PDN power integrity prediction, many of which still suffer from large training overhead, inefficiency, or non-scalability. Thus, this article proposed an efficient and scalable framework for the worst-case power integrity prediction, which can handle general tasks including dynamic noise prediction and bump current prediction. The framework first reduces the spatial and temporal redundancy in the PDN and input current vector and then employs efficient feature extraction as well as a novel convolutional neural network architecture to predict the worst-case power integrity. Experimental results show that the proposed framework consistently outperforms the commercial tool and the state-of-the-art machine learning method with only 0.63–1.02% mean relative error and 25–69× speedup for noise prediction and 0.22–1.06% mean relative error and 24–64× speedup for bump current prediction. Yufei Chen 0007, Yucheng Wang 0005, Tianming Ni, Zhiguo Shi 0001, Xunzhao Yin, Cheng Zhuo |
ACM Trans. Design Autom. Electr. Syst. | 8 |
| 2022 | Energy efficient data search design and optimization based on a compact ferroelectric FET content addressable memoryabstractContent Addressable Memory (CAM) is widely used for associative search tasks in advanced machine learning models and data-intensive applications due to the highly parallel pattern matching capability. Most state-of-the-art CAM designs focus on reducing the CAM cell area by exploiting the nonvolatile memories (NVMs). There exists only little research on optimizing the design and energy efficiency of NVM based CAMs for practical deployment in edge devices and AI hardware. In this paper, we propose a general compact and energy efficient CAM design scheme that alleviates the design overhead by employing just one NVM device in the cell. We also propose an adaptive matchline (ML) precharge and discharge scheme that further optimizes the search energy by fully reducing the ML voltage swing. We consider Ferroelectric field effect transistors (FeFETs) as the representative NVM, and present a 2T-1FeFET CAM array including a sense amplifier implementing the proposed ML scheme. Evaluation results suggest that our proposed 2T-1FeFET CAM design achieves 6.64×/4.74×/9.14×/3.02× better energy efficiency compared with CMOS/ReRAM/STT-MRAM/2FeFET CAM arrays. Benchmarking results show that our approach provides 3.3×/2.1× energy-delay product improvement over the 2T-2R/2FeFET CAM in accelerating query processing applications. Jiahao Cai, Mohsen Imani, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin |
DAC | 8 |
| 2022 | Worst-case dynamic power distribution network noise prediction using convolutional neural networkabstractWorst-case dynamic PDN noise analysis is an essential step in PDN sign-off to ensure the performance and reliability of chips. However, with the growing PDN size and increasing scenarios to be validated, it becomes very time- and resource-consuming to conduct full-stack PDN simulation to check the worst-case noise for different test vectors. Recently, various works have proposed machine learning based methods for supply noise prediction, many of which still suffer from large training overhead, inefficiency, or non-scalability. Thus, this paper proposed an efficient and scalable framework for the worst-case dynamic PDN noise prediction. The framework first reduces the spatial and temporal redundancy in the PDN and input current vector, and then employs efficient feature extraction as well as a novel convolutional neural network architecture to predict the worst-case dynamic PDN noise. Experimental results show that the proposed framework consistently outperforms the commercial tool and the state-of-the-art machine learning method with only 0.63--1.02% mean relative error and 25--69× speedup. Yufei Chen 0007, Xunzhao Yin, Cheng Zhuo |
DAC | 3 |
| 2022 | iMARS: an in-memory-computing architecture for recommendation systemsabstractRecommendation systems (RecSys) suggest items to users by predicting their preferences based on historical data. Typical RecSys handle large embedding tables and many embedding table related operations. The memory size and bandwidth of the conventional computer architecture restrict the performance of RecSys. This work proposes an in-memory-computing (IMC) architecture (iMARS) for accelerating the filtering and ranking stages of deep neural network-based RecSys. iMARS leverages IMC-friendly embedding tables implemented inside a ferroelectric FET based IMC fabric. Circuit-level and system-level evaluation show that iMARS achieves 16.8x (713x) end-to-end latency (energy) improvement compared to the GPU counterpart for the MovieLens dataset. Mengyuan Li 0001, Ann Franchesca Laguna, Dayane Reis, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
DAC | 4 |
| 2022 | HDPG: hyperdimensional policy-based reinforcement learning for continuous controlabstractTraditional robot control or more general continuous control tasks often rely on carefully hand-crafted classic control methods. These models often lack the self-learning adaptability and intelligence to achieve human-level control. On the other hand, recent advancements in Reinforcement Learning (RL) present algorithms that have the capability of human-like learning. The integration of Deep Neural Networks (DNN) and RL thereby enables autonomous learning in robot control tasks. However, DNN-based RL brings both high-quality learning and high computation cost, which is no longer ideal for currently fast-growing edge computing scenarios. Yang Ni 0001, Mariam Issa, Danny Abraham, Mahdi Imani, Xunzhao Yin, Mohsen Imani |
DAC | 5 |
| 2022 | Eva-CAM: A Circuit/Architecture-Level Evaluation Tool for General Content Addressable MemoriesabstractContent addressable memories (CAMs), a special-purpose in-memory computing (IMC) unit, support parallel searches directly in memory. There are growing interests in CAMs for data-intensive applications such as machine learning and bioinformatics. The design space for CAMs is rapidly expanding. In addition to traditional ternary CAMs (TCAMs), analog CAM (ACAM) and multi-bit CAM (MCAM) designs based on various non-volatile memory (NVM) devices have been recently introduced and may offer higher density, better energy efficiency, and non-volatility. Furthermore, aside from the widely-used exact match based search, CAM-based approximate matches have been proposed to further extend the utility of CAMs to new application spaces. For this memory architecture, evaluating different CAM design options for a given application is becoming more challenging. This paper presents Eva-CAM, a circuit/architecture-level modeling and evaluation tool for CAMs. Eva-CAM supports TCAM, ACAM, and MCAM designs implemented in non-volatile memories, for both exact and approximate match types. It also allows for the exploration of CAM array structures and sensing circuits. Eva-CAM has been validated with HSPICE simulation results and chip measurements. A comprehensive case study is described for FeFET CAM design space exploration. Liu Liu 0023, Mohammad Mehdi Sharifi, Ramin Rajaei, Arman Kazemi, Kai Ni 0004, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 6 |
| 2022 | Energy-Efficient Brain-Inspired Hyperdimensional Computing Using Voltage ScalingabstractRecently, brain-inspired hyperdimensional computing (HDC) has demonstrated promising capability in a wide range of applications such as medical diagnosis, human activity recognition, and voice classification, etc. Despite the growing popularity of HDC, its memory-centric computing characteristics make the associative memory implementation under significant energy consumption due to the massive data storage and processing. In this paper, we present a systematic case study to leverage the application-level error resilience of HDC to reduce the energy consumption of HDC associative memory by using voltage scaling. Evaluation results on various applications show that our proposed approach can achieve 47.6% energy saving on associative memory with a 1% accuracy loss. We further explore two low-cost error masking methods: word masking and bit masking, to mitigate the impact of voltage scaling-induced errors. Experimental results show that the proposed word masking (bit masking) method can further enhance energy saving up to 62.3% (72.5%) with accuracy loss ≤1%. Sizhe Zhang, Dongning Ma, Jeff Zhang 0001, Xunzhao Yin, Xun Jiao 0002 |
DATE | 5 |
| 2022 | COSIME: FeFET Based Associative Memory for In-Memory Cosine Similarity SearchabstractIn a number of machine learning models, an input query is searched across the trained class vectors to find the closest feature class vector in cosine similarity metric. However, performing the cosine similarities between the vectors in Von-Neumann machines involves a large number of multiplications, Euclidean normalizations and division operations, thus incurring heavy hardware energy and latency overheads. Moreover, due to the memory wall problem that presents in the conventional architecture, frequent cosine similarity-based searches (CSSs) over the class vectors requires a lot of data movements, limiting the throughput and efficiency of the system. To overcome the aforementioned challenges, this paper introduces COSIME, a general in-memory associative memory (AM) engine based on the ferroelectric FET (FeFET) device for efficient CSS. By leveraging the one-transistor AND gate function of FeFET devices, current-based translinear analog circuit and winner-take-all (WTA) circuitry, COSIME can realize parallel in-memory CSS across all the entries in a memory block, and output the closest word to the input query in cosine similarity metric. Evaluation results at the array level suggest that the proposed COSIME design achieves 333× and 90.5× latency and energy improvements, respectively, and realizes better classification accuracy when compared with an AM design implementing approximated CSS. The proposed in-memory computing fabric is evaluated for an HDC problem, showcasing that COSIME can achieve on average 47.1× and 98.5× speedup and energy efficiency improvements compared with an GPU implementation. Che-Kai Liu, Haobang Chen, Mohsen Imani, Kai Ni 0004, Arman Kazemi, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu, Liang Zhao 0004, Cheng Zhuo, Xunzhao Yin |
ICCAD | 11 |
| 2022 | Aging Aware Retraining for Memristor-based Neuromorphic ComputingabstractMemristor-based crossbars, which can achieve 1-2 orders of magnitude energy efficiency improvement over digital machines, have been introduced to accelerate the neural networks of machine learning tasks. Due to the high voltage pulses repeatedly applied onto memristors during programming and online tuning, the effective resistance ranges of the memristors actually decrease as a result of aging, which eventually impair the inference accuracy of the neural network running on the memristor-based crossbar. In this paper, we propose an algorithm-hardware co-design framework combining aging aware retraining and gradient sparsification to mitigate the impact of aging and extend the lifetime of the crossbar. Experimental results show that the proposed method can effectively increase the inference accuracy by up to 16% even with severe aging, while the crossbar lifetime can be extended by up to $2.7\times$. Wenwen Ye, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin |
ISCAS | 6 |
| 2022 | PAM: A Piecewise-Linearly-Approximated Floating-Point Multiplier With Unbiasedness and ConfigurabilityabstractApproximate computing is a promising alternative to improve energy efficiency for IoT devices on the edge. This work proposes a piecewise-linearly-approximated and unbiased floating-point approximate multiplier with run-time configurability. We provide a theoretically sound formulation that turns multiplication approximation to an optimization problem. With the formulation and findings, a multi-level architecture is proposed to easily incorporate run-time configurability and module execution parallelism. Finally, the proposed multiplier is further optimized to reduce the circuit implementation complexity, making the multiplier linearly dependent on the precision requirement, instead of quadratically or exponentially as in prior work. When compared to the prior state-of-the-art approximate floating-point multiplier, ApproxLP M. Imaniet al, “ApproxLP: Approximate multiplication with linearization and iterative error control,” inProc. ACM/IEEE Des. Autom. Conf., 2019, pp. 1–6., the proposed multiplier outperforms in all the aspects including accuracy, area, and delay. By replacing a full-precision floating-point multiplier in GPU, the proposed design can improve the energy efficiency for various edge computing tasks. Even with Level 1 approximation, the proposed multiplier improves energy efficiency up to 20× for machine learning on CIFAR-10, with almost negligible accuracy loss. Chuangtao Chen 0001, Weikang Qian, Mohsen Imani, Xunzhao Yin, Cheng Zhuo |
IEEE Trans. Computers | 4 |
| 2022 | FeFET Multi-Bit Content-Addressable Memories for In-Memory Nearest Neighbor SearchabstractNearest neighbor (NN) search computations are at the core of many applications such as few-shot learning, classification, and hyperdimensional computing. As such, efficient hardware support for NN search is highly desired. In-memory computing using emerging devices offers attractive solutions for NN search. Solutions based on ternary content-addressable memories (TCAMs) offer high energy and latency improvements for NN search at the expense of accuracy. In this work, we propose a novel distance function that can be natively evaluated with multi-bit content-addressable memories (MCAMs) based on ferroelectric FETs (FeFETs) to perform a single-step, in-memory NN search. We evaluate the efficacy of FeFET MCAMs in the context of few-shot learning applications with different datasets. As an example, we achieve a 78.54% accuracy for a 5-way, 5-shot classification task for the mini-ImageNet dataset (only 1.5% lower than software-based implementations) when using a 3-bit MCAM for NN search. We consider the effects of FeFET threshold voltage variations on the application accuracy and analyze the area and search energy requirements of FeFET MCAMs for accurate operations. Our results indicate that MCAMs require 2× lower area and search energy than TCAMs to achieve the same accuracy. Furthermore, we experimentally demonstrate a 2-bit implementation of FeFET MCAM using AND arrays from GLOBALFOUNDRIES to further validate the design concept. Arman Kazemi, Mohammad Mehdi Sharifi, Ann Franchesca Laguna, Franz Müller 0001, Xunzhao Yin, Thomas Kämpfe, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Computers | 5 |
| 2022 | Improving Fault Tolerance for Reliable DNN Using Boundary-Aware ActivationabstractIn this article, we approach to construct reliable deep neural networks (DNNs) for safety-critical artificial intelligent applications. We propose to modify rectified linear unit (ReLU), a commonly used activation function in DNNs, to tolerate the faults incurred by bit-flip perturbation on weights. Through theoretic analysis of the fault propagation in the layers with ReLU activation, we observe that bounding the output of ReLU activation can help to tolerate the weight faults. Then, we propose a novel ReLU design called boundary-aware ReLU (BReLU) to improve the reliability of DNNs, in which an upper bound of ReLU is determined such that the deviation between the boundary and original outputs cannot affect the final result. We propose a gradient-ascent-based algorithm to find the boundaries for BReLU activations of all DNN layers. Without retraining the network, our approach is cost effective and practical when deployed in safety-critical artificial intelligent systems. Detailed experiments and real-life application benchmarking demonstrate that our approach can improve the accuracy of DNN VGG16 from 16.7% to 82.6% on average assuming the practical weight faults, with only 13% memory and 2.78% time overhead, respectively. Jinyu Zhan, Ruoxu Sun, Wei Jiang 0016, Yucheng Jiang, Xunzhao Yin, Cheng Zhuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2022 | VirtualSync+: Timing Optimization With Virtual SynchronizationabstractIn digital circuit designs, sequential components such as flip-flops are used to synchronize signal propagations. Logic computations are aligned at and thus isolated by flip-flop stages. Although this fully synchronous style can reduce design efforts significantly, it may affect circuit performance negatively, because sequential components can only introduce delays into signal propagations but never accelerate them. In this article, we propose a new timing model, VirtualSync+, in which signals, specially those along critical paths, are allowed to propagate through several sequential stages without flip-flops. Timing constraints are still satisfied at the boundary of the optimized circuit to maintain a consistent interface with existing designs. By removing clock-to-q delays and setup time requirements of flip-flops on critical paths, the performance of a circuit can be pushed even beyond the limit of traditional sequential designs. In addition, we further enhance the optimization with VirtualSync+ by fine-tuning with commercial design tools, e.g., design compiler from Synopsys, to achieve more accurate result. To achieve this fine-tuning, we first optimize the circuits by reallocating sequential components with sequential and combinational components as delay units. Afterward, the removal locations of flip-flops with respect to the circuits under optimization are extracted and the corresponding wave-pipelining timing constraints compatible with commercial design tools are established. These timing constraints are then incorporated into the optimization flow of commercial tools to generate the optimized circuits. The experimental results demonstrate that circuit performance can be improved by up to 4% (average 1.5%) compared with that after extreme retiming and sizing, while the increase of area is still negligible. This timing performance is enhanced beyond the limit of traditional sequential designs. It also demonstrates that compared with those after retiming and sizing, the circuits with VirtualSync+ can achieve better timing performance under the same area cost or smaller area cost under the same clock period, respectively. Grace Li Zhang, Bing Li 0005, Xing Huang 0001, Xunzhao Yin, Cheng Zhuo, Masanori Hashimoto, Ulf Schlichtmann |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2022 | Senputing: An Ultra-Low-Power Always-On Vision Perception Chip Featuring the Deep Fusion of Sensing and ComputingabstractAlways-on intelligent visual perception applications are widely deployed in edges in the AIoT era. In order to eliminate power costs of data conversion and transmission, this paper proposes Senputing, an ultra-low-power processing-in-sensor chip that completely fuses sensing and computing together for a BNN-based hierarchical processing system. This chip could operate in two modes. In computation mode, photocurrents are directly utilized for computing without being converted into voltages, and the computation results of 1-st BNN layer are directly sent out to subsequent BNN processors for an always-on coarse classification, eliminating conversion power and storage cost of raw images. Once an interested objected is detected, this chip switches to sensor mode and sends raw images to potential full-precision processors or cloud servers for fine-grained recognition or segmentation. A$32\times 32$prototype is fabricated with 180nm CMOS process. It accomplishes MNIST dataset classification task with the accuracy of 93.76% and the power consumption of 147nW at 156fps, achieving$13.1\times $energy efficiency compared with state-of-the-art work. Han Xu 0006, Ningchao Lin, Qi Wei 0001, Runsheng Wang, Cheng Zhuo, Xunzhao Yin, Fei Qiao, Huazhong Yang |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2021 | Cross-layer Design for Computing-in-Memory: From Devices, Circuits, to Architectures and ApplicationsabstractThe era of Big Data, Artificial Intelligence (AI) and Internet of Things (IoT) is approaching, but our underlying computing infrastructures are not sufficiently ready. The end of Moore's law and process scaling as well as the memory wall associated with von Neumann architectures have throttled the rapid development of conventional architectures based on CMOS technology, and cross-layer efforts that involve the interactions from low-end devices to high-end applications have been prominently studied to overcome the aforementioned challenges. On one hand, various emerging devices, e.g., Ferroelectric FET, have been proposed to either sustain the scaling trends or enable novel circuit and architecture innovations. On the other hand, novel computing architectures/algorithms, e.g., computing-in-memory (CiM), have been proposed to address the challenges faced by conventional von Neumann architectures. Naturally, integrated approaches across the emerging devices and computing architectures/algorithms for data-intensive applications are of great interests. This paper uses the FeFET as a representative device, and discuss about the challenges, opportunities and contributions for the emerging trends of cross-layer co-design for CiM. Hussam Amrouch, Xiaobo Sharon Hu, Mohsen Imani, Ann Franchesca Laguna, Michael T. Niemier, Simon Thomann, Xunzhao Yin, Cheng Zhuo |
ASP-DAC | 7 |
| 2021 | Robustness of Neuromorphic Computing with RRAM-based Crossbars and Optical Neural NetworksabstractRRAM-based crossbars and optical neural networks are attractive platforms to accelerate neuromorphic computing. However, both accelerators suffer from hardware uncertainties such as process variations. These uncertainty issues left unaddressed, the inference accuracy of these computing platforms can degrade significantly. In this paper, a statistical training method where weights under process variations and noise are modeled as statistical random variables is presented. To incorporate these statistical weights into training, the computations in neural networks are modified accordingly. For optical neural networks, we modify the cost function during software training to reduce the effects of process variations and thermal imbalance. In addition, the residual effects of process variations are extracted and calibrated in hardware test, and thermal variations on devices are also compensated in advance. Simulation results demonstrate that the inference accuracy can be improved significantly under hardware uncertainties for both platforms. Grace Li Zhang, Bing Li 0005, Ying Zhu 0008, Yiyu Shi 0001, Xunzhao Yin, Cheng Zhuo, Huaxi Gu, Tsung-Yi Ho, Ulf Schlichtmann |
ASP-DAC | 6 |
| 2021 | Bayesian Inference Based Robust Computing on Memristor CrossbarabstractMemristor based crossbars are a promising platform for neural network acceleration. To deploy a trained network model on a memristor crossbar, memristors need to be programmed to realize the trained weights of the network. However, due to process and dynamic variations, deviation of weights from the trained value is inevitable and inference accuracy thus degrades. In this paper, we propose a unified Bayesian inference based framework which connects hardware variations and algorithmic training together for robust computing on memristor crossbars. The framework incorporates different levels of variations into priori weight distribution, and transforms robustness optimization to Bayesian neural network training, where weights of neural networks are optimized to accommodate variations and minimize inference degradation. Simulation results with the proposed framework confirm stable inference accuracy under process and dynamic variations. Qingrong Huang, Grace Li Zhang, Xunzhao Yin, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo |
DAC | 4 |
| 2021 | RegHD: Robust and Efficient Regression in Hyper-Dimensional Learning SystemabstractMachine learning (ML) algorithms are key enablers to effectively assimilate and extract information from many generated data in the Internet of Things. However, running ML algorithms often results in extremely slow processing speed and high energy consumption. To achieve real-time performance with high energy efficiency and robustness, we proposed RegHD, the first regression solution based on Hyperdimensional computing. RegHD redesign a regression algorithm using strategies that more closely model the ultimate efficient learning machine: the human brain. RegHD performs regression after mapping data points into high-dimensional space using similarity preserving encoding. Due to the encoder’s non-linearity, RegHD learns a regression model in an efficient and linear way. RegHD creates two set of models: Input Model to cluster data points with high similarity, and Regression Model to generate a regression model for each clustered data. During prediction, RegHD computes the output value by the weighted accumulation of all regression models, considering the model confidence obtained during similarity search. To improve RegHD efficiency, we also proposed a framework that enables RegHD model quantization while having no impact on the learning accuracy. Our evaluation shows that RegHD provides 5.6 × and 12.3 × (2.9 × and 4.2 ×) faster and energy efficient training (inference) as compared to state-of-the-art regression algorithms, while providing similar quality of learning. Alejandro Hernández-Cano, Cheng Zhuo, Xunzhao Yin, Mohsen Imani |
DAC | 3 |
| 2021 | Cognitive Correlative Encoding for Genome Sequence Matching in Hyperdimensional SystemabstractPattern matching is one of the key algorithms in identifying and analyzing genomic data. In this paper, we propose HYPERS, a novel framework supporting highly efficient and parallel pattern matching based on HyperDimensional computing (HDC). HYPERS transforms inherent sequential processes of pattern matching to highly-parallelizable computation tasks using HDC. HYPERS exploits HDC memorization to encode and represent the genome sequences using high-dimensional vectors. Then, it combines the genome sequences to generate an HDC reference library. During the matching, HYPERS performs alignment by exact or approximate similarity check of an encoded query with the HDC reference library. HYPERS functionality is supported by theoretical proof, verified by software implementation, and extensively tested on the existing hardware platform. Our evaluation on FPGA shows that HYPERS provides, on average, $ 17.5\times$ speedup and $ 39.4\times$ energy efficiency as compared to the state-of-the-art pattern matching tools running on GTX 1080 GPU. Prathyush Poduval, Zhuowen Zou, Xunzhao Yin, Elaheh Sadredini, Mohsen Imani |
DAC | 3 |
| 2021 | Joint Sparsity with Mixed Granularity for Efficient GPU Implementation
Chuliang Guo, Xingang Yan, Yufei Chen 0007, He Li 0008, Xunzhao Yin, Cheng Zhuo |
DATE | 5 |
| 2021 | Energy-Aware Designs of Ferroelectric Ternary Content Addressable MemoryabstractTernary content addressable memories (TCAMs) are a special form of computing-in-memory (CiM) circuits that aim to address the so-called memory wall issues by merging the parallel search function with memory blocks. Due to the content addressing nature, TCAMs have been widely utilized for search intensive tasks in low-power, data analytic applications, such as IP routers, associative memories, and learning models. While most state-of-the-art TCAM designs focus on improving the TCAM density by harnessing compact nonvolatile memories (NVMs), little efforts have been spent on reducing and optimizing the energy consumption of the NVM based TCAM. In this paper, by exploiting the Ferroelectric FET (FeFET) as a representative NVM, we propose two compact and energy-aware designs of ferroelectric TCAMs for low power applications. We first introduce a novel 2FeFET based XOR-like gate structure that can also be adopted to other NVMs, and then leverage the structure to propose two TCAM designs that achieve high energy efficiency by either reducing the associated precharge overhead (2FeFET-1T cell), or eliminating the precharge phase typically required by TCAMs (2FeFET-2T cell). We evaluate and compare the designs w.r.t area, search energy and delay at array level with other existing designs, and benchmark the proposed TCAM designs in an associative memory based GPU architecture. The results suggest that the proposed 2FeFET-1T/2FeFET-2T TCAM design consumes 3.03X/8.08X less search energy than the conventional 16T CMOS TCAM, while the proposed design cell area is only 32.1%/39.3% of the latter. Compared with the state-of-the-art 2FeFET only TCAM array, our proposed designs still achieve 1.79X and 4.79X search energy reduction, respectively. Moreover, our proposed designs can achieve, on average, 45.2%/51.5% energy saving compared with the conventional GPU based architecture at the application level. Yu Qian 0002, Zhenhao Fan, Chao Li 0065, Mohsen Imani, Kai Ni 0004, Grace Li Zhang, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo, Xunzhao Yin |
DATE | 11 |
| 2021 | Real-Time and Robust Hyperdimensional ClassificationabstractHyper-Dimensional computing (HDC) is a brain-inspired learning approach for efficient and robust learning on today's embedded devices. HDC supports single-pass learning, where it generates a classification model by one-time looking at each training data point. However, the single-pass model provides weak classification accuracy due to model saturation caused by naively accumulating high-dimensional data. Although the retraining model for hundreds of iterations addresses the model saturation and boosts the accuracy, it comes with significant training costs. In this paper, we propose OnlineHD, an adaptive HDC training framework for accurate, efficient, and robust learning. During single-pass training, OnlineHD identifies common patterns and eliminates model saturation. For each data point, OnlineHD updates the model depending on how similar it is to the existing model, instead of naive data accumulation. We expand the OnlineHD framework to support highly-accurate iterative training. We also exploit the holographic distribution of patterns in high-dimensional space to make OnlineHD ultra-robust against possible noise and hardware failure. Our evaluations on a wide range of classification problems show that OnlineHD adaptive training provides comparable classification accuracy to the retrained model while getting all efficiency benefits that a singlepass training provides. Alejandro Hernández-Cano, Cheng Zhuo, Xunzhao Yin, Mohsen Imani |
ACM Great Lakes Symposium on VLSI | 3 |
| 2021 | ICCAD Tutorial Session Paper Ferroelectric FET Technology and Applications: From Devices to SystemsabstractThe rapidly increasing volume and complexity of data is demanding the relentless scaling of computing power. With transistor feature size approaching physical limits, the benefits that CMOS technology can provide is diminishing. For future energy efficient computing systems, researchers aim to exploit various emerging nanotechnologies to replace conventional CMOS technology. In particular, ferroelectric FETs (FeFETs) appear to be a promising candidate to continue improving energy efficiency for data-intensive applications. Advances in FeFET scalability and FeFET compatibility with CMOS have sparked growing interest in device, circuit, and system communities. While FeFET is still evolving, many researchers and developers are already cautiously optimistic about its future. This paper provides a review on FeFET's recent technology advances, challenges, and opportunities, with a particular emphasis upon device modeling and circuit design of FeFET content addressable memory, as well as their applications in machine learning. Hussam Amrouch, Xiaobo Sharon Hu, Arman Kazemi, Ann Franchesca Laguna, Kai Ni 0004, Michael T. Niemier, Mohammad Mehdi Sharifi, Simon Thomann, Xunzhao Yin, Cheng Zhuo |
ICCAD | 10 |
| 2021 | Reliable Memristor-based Neuromorphic Design Using Variation- and Defect-Aware TrainingabstractThe memristor crossbar provides a unique opportunity to develop a neuromorphic computing system (NCS) with high scalability and energy efficiency. However, the reliability issues that arise from the immature fabrication process and physical device limitations, i.e., variations and stuck-at-faults (SAF), dramatically prevent its wide application in practice. Specifically, variations make the programmed weights deviate from their expected values. On the other hand, defective mem-ristors cannot even represent the weights effectively. In this work, we propose a variation- and defect-aware framework to improve the reliability of memristor-based NCS while minimizing the inference performance loss. We propose to develop analytical weight models to characterize the non-ideal effects of variations and SAFs, which can then be incorporated into a Bayesian neural network as priori and constraint. We then convert the reliability improvement to the neural network training for optimal weights that can accommodate variations and defects across the chips, which does not require computation-intensive retraining or cost-expensive testing. Extensive experimental results with the proposed framework confirm its effective capability of improving the reliability of NCS, while significantly mitigating the inference accuracy degradation under even severe variations and SAFs. Grace Li Zhang, Xunzhao Yin, Bing Li 0005, Ulf Schlichtmann, Cheng Zhuo |
ICCAD | 3 |
| 2021 | On the Reliability of In-Memory Computing: Impact of Temperature on Ferroelectric TCAMabstractWith the rapid development of emerging technologies, especially the ferroelectric field-effect transistors (FeFETs), the density and energy efficiency of ternary content addressable memory (TCAM) have been increasingly improved. TCAM plays a major role in realizing In-Memory Computing and other brain-inspired computing concepts. Recently, the parallel search functionality of a FeFET based ultra-dense TCAM design is also enhanced with a Hamming distance-based approximate search scheme. However, in order to realize the highly-promising TCAM design, in which the approximate search function based on Hamming distance is implemented, it is inevitable to investigate the impact of temperature on the reliability of FeFET-based TCAM cells as well as all involved peripheral circuits. In this paper, the temperature impact on the FeFET at the device level and the approximate TCAM design at the circuit level is investigated for the first time. The demonstrated example of a FeFET-based TCAM array shows that the unique temperature dependency of a FeFET device can help mitigate the temperature impact on the FeFET TCAM array. Based on the observation, we showcase, evaluate, and discuss in detail one strategy to eliminate the temperature impact on the approximate TCAM design. Understanding and mitigating the deleterious impact of temperature on the reliability of FeFET-based TCAM circuits is essential to ensure reliable In-Memory Computing. Simon Thomann, Chao Li 0065, Cheng Zhuo, Om Prakash 0007, Xunzhao Yin, Xiaobo Sharon Hu, Hussam Amrouch |
VTS | 5 |
| 2021 | A Reconfigurable Multiplier for Signed Multiplications with Asymmetric Bit-WidthsabstractMultiplications have been commonly conducted in quantized CNNs, filters, and reconfigurable cores, and so on, which are widely deployed in mobile and embedded applications. Most multipliers are designed to perform multiplications with symmetric bit-widths, i.e., n - by n -bit multiplication. Such features would cause extra area overhead and performance loss when m - by n -bit multiplications ( m > n ) are deployed in the same hardware design, resulting in inefficient multiplication operations. It is highly desired and challenging to propose a reconfigurable multiplier design to accommodate operands with both symmetric and asymmetric bit-widths. In this work, we propose a reconfigurable approximate multiplier to support multiplications at various precisions, i.e., bit-widths. Unlike prior works of approximate adders assuming a uniform weight distribution with bit-wise independence, scenarios like a quantized CNN may have a centralized weight distribution and hence follow a Gaussian-like distribution with correlated adjacent bits. Thus, a new block-based approximate adder is also proposed as part of the multiplier to ensure energy-efficient operation with an awareness of the bit-wise correlation. Our experimental results show that the proposed approximate adder significantly reduces the error rate by 76% to 98% over a state-of-the-art approximate adder for Gaussian-like distribution scenarios. Evaluation results show that the proposed multiplier is 19% faster and 22% more power saving than a Xilinx multiplier IP at the same bit precision and achieves a 23.94-dB peak signal-to-noise ratio, which is comparable to the accurate one of 24.10 dB when deployed in a Gaussian filter for image processing tasks. Chuliang Guo, Li Zhang 0021, Grace Li Zhang, Bing Li 0005, Weikang Qian, Xunzhao Yin, Cheng Zhuo |
ACM J. Emerg. Technol. Comput. Syst. | 7 |
| 2020 | Nonvolatile and Energy-Efficient FeFET-Based Multiplier for Energy-Harvesting DevicesabstractEnergy-harvesting internet-of-things devices must deal with unstable power input. Nonvolatile processors (NVPs) can offer an effective solution. Compact and low-energy arithmetic circuits that can efficiently switch between computation and backup operations are highly desirable for NVP design. This paper introduces a nonvolatile ferroelectric field-effect transistors (FeFET)-based sequential multiplier with the ability to do continued calculation after a power outage, thus achieving zero backup overhead. We exploit the unique characteristics of FeFETs to construct key components of a sequential multiplier. The multiplier relies on a FeFET-based adder and a new FeFET-based latch to achieve compact area and low operating energy. Moreover, it uses the hysteretic characteristic of FeFETs to realize the storage capability, and hence is able to store, at no extra cost, the intermediate data of an operation in a nonvolatile manner. This property provides support for continued computation when power supplies may be intermittent. Simulation results show that, assuming the same technology node, the proposed FeFET-based multiplier saves up to 21% and 19% area than a conventional CMOS-based sequential multiplier of 4-bits and 8-bits, respectively. It also saves 32% and 73% less area compared with a CMOS-based array multiplier. Furthermore, the proposed design can offer up to 32%/23% energy saving per operation compared with a 4/8-bit CMOS-based sequential multiplier. Mengyuan Li 0001, Xunzhao Yin, Xiaobo Sharon Hu, Cheng Zhuo |
ASP-DAC | 2 |
| 2020 | Emerging Neural Workloads and Their Impact on HardwareabstractWe consider existing and emerging neural workloads, and what hardware accelerators might be best suited for said workloads. We begin with a discussion of analog crossbar arrays, which are known to be well-suited for matrix-vector multiplication operations that are commonplace in existing neural network models such as convolutional neural networks (CNNs). We highlight candidate crosspoint devices, what device and materials challenges must be overcome for a given device to be employed in a crossbar array for a computationally interesting neural workload, and how circuit and algorithmic optimizations may be employed to mitigate undesirable characteristics from devices/materials. We then discuss two emerging neural workloads. We first consider machine learning models for one- and few-shot learning tasks (i.e., where a network can be trained with just one or a few, representative examples of a given class). Notably crossbar-based architectures can be used to accelerate said models. Hardware solutions based on content addressable memory arrays will also be discussed. We then consider machine learning models for recommendation systems. Recommendation models, an emerging class of machine learning models, employ distinct neural network architectures that operate of continuous and categorical input features which make hardware acceleration challenging. We will discuss the open research challenges and opportunities within this space. David Brooks 0001, Martin M. Frank, Tayfun Gokmen, Udit Gupta 0001, Xiaobo Sharon Hu, Shubham Jain 0004, Ann Franchesca Laguna, Michael T. Niemier, Ian O'Connor, Anand Raghunathan, Ashish Ranjan 0001, Dayane Reis, Jacob R. Stevens, Carole-Jean Wu, Xunzhao Yin |
DATE | 15 |
| 2020 | AxR-NN: Approximate Computation Reuse for Energy-Efficient Convolutional Neural NetworksabstractThe recent success of convolutional neural networks (CNN) has led its implementation in specialized accelerators such as graphics processing unit (GPUs). However, the intensive computing workloads of CNNs remain a challenge to existing accelerators. By leveraging the error tolerance of CNNs, we propose a novel method to design energy-efficient CNN accelerators using approximate computation reuse (ACR), referred to as AxRNN. Computation reuse aims to reuse the previously computed results to avoid redundant executions. However, it cannot be applied directly to CNNs because CNNs do not have enough data locality. Thus, AxRNN performs approximate computation reuse under relaxed precision requirements on input patterns and design a reconfigurable architecture to support the ACR. This reconfigurable pattern matching is central to achieve a "controllable approximation". We implement the AxRNN using content addressable memory and integrate them with floating point units. Simulation results show that AxRNN reduces the computation energy by 30-58% with only 1-2.5% accuracy degradation on MNIST, EMNIST, and CIFAR-10 dataset. Dongning Ma, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu, Xun Jiao 0002 |
ACM Great Lakes Symposium on VLSI | 2 |
| 2020 | Modeling and Benchmarking Computing-in-Memory for Design Space ExplorationabstractThe bottleneck between the limited memory bandwidth and high speed processing demands is the main cause of problems associated with high volume of data transfers in data-intensive applications. As a possible remedy to these issues, computing-in-memory (CiM) enables a subset of logic and arithmetic operations to be performed where the data resides, i.e., inside the memory. Various CiM designs have been proposed to date, based on different technologies. Given the variety of options available, picking the right design option for a system/application can be a complex task. When choosing a CiM design, it is important to establish evaluation conditions that are as uniform as possible to make a fair choice between available design options. In this paper, we describe a methodology for an uniform benchmarking of CiM designs. Our approach evaluates devices/circuits, arrays and the overall impact of CiM to a system with a framework based on Eva-CiM. As a case study, we analyze the array-level performance of 7 recent CiM designs implemented with SRAM, DRAM, FeFET-RAM, STT-MRAM, SOT-MRAM, and RRAM. After we identify that the FeFET-RAM-based design shows promising energy and delay savings at the array level, we carry out a system level evaluation showing that FeFET-RAM-based CiM outperforms a CMOS SRAM CiM baseline by an average of 60% across a set of 17 benchmarks (with respect to energy savings). Regarding speedups, both technologies offer virtually the same benefit of about 1.5X when compared to a situation where processing does not happen in memory. Dayane Reis, Shaahin Angizi, Xunzhao Yin, Deliang Fan, Michael T. Niemier, Cheng Zhuo, Xiaobo Sharon Hu |
ACM Great Lakes Symposium on VLSI | 4 |
| 2020 | Optimally Approximated and Unbiased Floating-Point Multiplier with Runtime ConfigurabilityabstractApproximate computing is a promising alternative to improve energy efficiency for IoT devices on the edge. This work proposes an optimally approximated and unbiased floating-point approximate multiplier with runtime configurability. We provide a theoretically sound formulation that turns multiplication approximation to an optimization problem. With the formulation and findings, a multilevel architecture is proposed to easily incorporate runtime configurability and module execution parallelism. Finally, an optimization scheme is applied to improve the area, making it linearly dependent on the precision, instead of quadratically or exponentially as in prior work. In addition to the optimal approximation and configurability, the proposed design has an efficient circuit implementation that uses inversion, shift and addition instead of complex arithmetic operations. When compared to the prior state-of-the-art approximate floating-point multiplier, ApproxLP [30], the proposed design outperforms in all aspects including accuracy, area, and delay. By replacing the regular full-precision multiplier in GPU, the proposed design can improve the energy efficiency for various edge computing tasks. Even with Level 1 approximation, the proposed design improves energy efficiency up to 122× for machine learning on CIFAR-10, with almost negligible accuracy loss. Chuangtao Chen 0001, Weikang Qian, Mohsen Imani, Xunzhao Yin, Cheng Zhuo |
ICCAD | 5 |
| 2020 | Seed-and-Vote based In-Memory Accelerator for DNA Read MappingabstractGenome analysis is becoming more important in the fields of forensic science, medicine, and history. Sequencing technologies such as High Throughput Sequencing (HTS) and Third Generation Sequencing (TGS) have greatly accelerated genome sequencing. However, genome read mapping remains significantly slower than sequencing. Because of the enormous amount of data needed, the speed of the data transfer between the memory and the processing unit limits the execution speed. In-memory computing can help address the memory-bandwidth bottleneck by minimizing data transfers. Ternary Content Addressable Memories (TCAMs) have been used in accelerators because of their fast searching capability for seed-and-extend, a popular read mapping approach. Seed-and-vote, another read mapping approach, is faster than the seed-and-extend approach but has lower accuracies when used with very short reads. Since sequencing technology is moving to longer reads, the seed-and-vote approach is becoming more viable. We propose a genome read mapping accelerator that uses approximate TCAM to execute the Fast Seed and Vote algorithm (FSVA) that can map both short and long reads. We achieved 400X acceleration compared to the seed-and-extend approach BWA-MEM on a CPU and 115X acceleration at 30X energy improvement compared to state-of-the-art in-memory accelerator using the seed-and-extend approach at 98.75% accuracy for 100bp reads. Ann Franchesca Laguna, Hasindu Gamaarachchi, Xunzhao Yin, Michael T. Niemier, Sri Parameswaran, Xiaobo Sharon Hu |
ICCAD | 3 |
| 2020 | Countering Variations and Thermal Effects for Accurate Optical Neural NetworksabstractOptical neural networks (ONNs) have emerged as a promising high-performance computing platform to accelerate deep neural networks. In ONNs, phases of light are modulated through Mach-Zehnder Interferometers (MZIs), and MZIs are connected in a gridlike layout to implement multiply-accumulate operations. However, ONNs are very sensitive to process variations and thermal effects. This sensitivity leads to a significant degradation of inference accuracy of ONNs and thus renders them unusable in practice. In this paper, we propose a framework to calibrate process variations and counter thermal effects by power compensation. Experimental results demonstrate that the proposed framework can recover the inference accuracy under variations and thermal effects, e.g., from as low as 11.05% back to 74.11% for LeNet-5 on Cifar10, so that ONNs can achieve an inference accuracy similar to the accuracy after software training while providing their high bandwidth in neuromorphic computing. Ying Zhu 0008, Grace Li Zhang, Bing Li 0005, Xunzhao Yin, Cheng Zhuo, Huaxi Gu, Tsung-Yi Ho, Ulf Schlichtmann |
ICCAD | 4 |
| 2020 | A Convolutional Neural Network Accelerator Architecture with Fine-Granular Mixed Precision ConfigurabilityabstractConvolutional neural networks (CNNs) have been widely deployed in deep learning applications, especially on power hungry GP-GPUs. Recent efforts in designing CNN accelerators are considered as a promising alternative to achieve higher energy efficiency. Unfortunately, with the growing complexity of CNN, the demanded computational and storage resources for accelerators keep increasing, hindering its wider applications in mobile devices. On the other hand, many quantization algorithms have been proposed for efficient CNN training, which brings many small or zero weights. This is a unique opportunity for accelerator designers to employ much fewer bits, e.g., 4 bits, in both arithmetic core and storage, thereby saving significant design cost. However, such a single precision strategy inevitably compromises the accuracy as some key operations may demand a higher precision. Thus, this paper proposes a low power CNN accelerator architecture that can simultaneously conduct computations with mixed precisions and assign the appropriate arithmetic cores to operation with different precision demands. This proposed architecture can achieve significant area and energy savings, without accuracy compromise. The experimental results show that the proposed architecture implemented on FPGA can reduces almost half of the weight storage and MAC area, and lower the dynamic power by 12.1% when compared with a state-of-the-art CNN accelerator design. Li Zhang 0021, Chuliang Guo, Xunzhao Yin, Cheng Zhuo |
ISCAS | 4 |
| 2020 | SearcHD: A Memory-Centric Hyperdimensional Computing With Stochastic TrainingabstractBrain-inspired hyperdimensional (HD) computing emulates cognitive tasks by computing with long binary vectors-also know as hypervectors-as opposed to computing with numbers. However, we observed that in order to provide acceptable classification accuracy on practical applications, HD algorithms need to be trained and tested on nonbinary hypervectors. In this article, we propose SearcHD, a fully binarized HD computing algorithm with a fully binary training. SearcHD maps every data points to a high-dimensional space with binary elements. Instead of training an HD model with nonbinary elements, SearcHD implements a full binary training method which generates multiple binary hypervectors for each class. We also use the analog characteristic of nonvolatile memories (NVMs) to perform all encoding, training, and inference computations in memory. We evaluate the efficiency and accuracy of SearcHD on a wide range of classification applications. Our evaluation shows that SearcHD can provide on average 31.1× higher energy efficiency and 12.8× faster training as compared to the state-of-the-art HD computing algorithms. Mohsen Imani, Xunzhao Yin, John Messerly, Saransh Gupta, Michael T. Niemier, Xiaobo Sharon Hu, Tajana Rosing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2019 | Ferroelectric FET Based In-Memory Computing for Few-Shot LearningabstractAs CMOS technology advances, the performance gap between the CPU and main memory has not improved. Furthermore, the hardware deployed for Internet of Things (IoT) applications need to process ever growing volumes of data, which can further exacerbate the "memory wall". Computing-in-memory (CiM) architectures, where logic and arithmetic operations are performed in memory, can significantly reduce energy and latency overheads associated with data transfer, and potentially alleviate processor-memory bottlenecks. In this paper, we consider the utility of ternary content addressable memory (TCAM) arrays and CiM arrays based on ferroelectric field effect transistors (FeFETs) to support emerging machine learning models that can learn new classes of data with significantly less training overhead - highly desirable in IoT applications. Architecturally, we use TCAM and CiM arrays to implement the external memory module in a memory enhanced neural network (MENN) - which can be used to minimize catastrophic forgetting - a major problem in applications such as lifelong and few-shot learning. As a representative example, we achieve 95.14% accuracy for a few-shot learning task with the Omniglot data set by using a combined L∞ infinity and L1 distance metric computed via a TCAM-CiM cascaded architecture (as opposed to 99.06% accuracy assuming a GPU backed by DRAM). While there is a slight drop in accuracy, the TCAM-CiM approach is 4.34X faster and 4.18X more energy efficient than a CMOS implementation for the same task. The ability of an FeFET to serve as both a compact logic and storage element helps to enable dense CiM and TCAM structures that drive the aforementioned improvements to application-level figures of merit (FOMs). Ann Franchesca Laguna, Xunzhao Yin, Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu |
ACM Great Lakes Symposium on VLSI | 2 |
| 2019 | The Impact of Emerging Technologies on Architectures and System-level Management: Invited PaperabstractThe goal of this work is to introduce and discuss different kinds of emerging technologies for logic circuitry and memory with respect to the key question of how they will impact future system-on-chip architectures and system-level management techniques. It is obvious that emerging technologies should have an impact there in order to fully exploit their technological advantages but also in order to deal with any disadvantages they might come with. In this special session paper, three promising emerging technologies are presented: (i) Negative Capacitance Field-Effect Transistor (NCFET) as a new CMOS technology with advantages primarily for low-power design, (ii) Ferroelectric FET (FeFET) as a non-volatile, area-efficient and low-power combined logic and memory as well as (iii) a Phase-Change Memory (PCM) and Resistive RAM (ReRAM) offering a large potential for tackling the memory wall problem in the von Neumann architecture. Our analysis demonstrates that not only new computing paradigms are promoted by these new technologies, it will also be seen that the trade-offs between the classical design parameters of low power, performance etc. will shift and hence emerging technologies will offer new Pareto points in the design space of future on-chip architectures. In that context, this work is unique as it bridges the gap between the technology side and system/architecture-level side to draw a vision of new technologies and their impact on architectures and system-level management. Jörg Henkel, Hussam Amrouch, Martin Rapp, Sami Salamin, Dayane Reis, Xunzhao Yin, Michael T. Niemier, Cheng Zhuo, Xiaobo Sharon Hu, Hsiang-Yun Cheng, Chia-Lin Yang |
ICCAD | 7 |
| 2019 | Ferroelectric FETs-Based Nonvolatile Logic-in-Memory CircuitsabstractAmong the beyond-complementary metal-oxide- semiconductor (CMOS) devices being explored, ferroelectric field-effect transistors (FeFETs) are considered as one of the most promising. FeFETs are being studied by all major semiconductor manufacturers, and experimentally, FeFETs are making rapid progress. FeFETs also stand out with the unique hysteretic Ids-Vgs characteristic that allows a device to function as both a switch and a nonvolatile (NV) storage element. We exploit this FeFET property to build two categories of fine-grained logic-in-memory (LiM) circuits: 1) ternary content addressable memory (TCAM) which integrates efficient and compact logic/processing elements into various levels of memory hierarchy; 2) basic logic function units for constructing larger and more complex LiM circuits. Two writing schemes (with and without negative supply voltages respectively) for FeFETs are introduced in our LiM designs. The resulting designs are compared with existing LiM approaches based on CMOS, magnetic tunnel junctions (MTJs), resistive random access memories (ReRAMs), ferrorelectric tunnel junctions (FTJs), etc., that afford the same circuit-level functionality. Simulation results show that FeFET-based NV TCAMs offer lower area overhead than MTJ (79%) and CMOS (42% less) equivalents, as well as better search energy-delay products (EDPs) than TCAM designs based on MTJ (149×), ReRAM (1.7×), and CMOS (1.3×) in array evaluations. NV FeFET-based LiM basic circuit blocks are also more efficient than functional equivalents based on MTJs in terms of propagation delay (4.2×) and dynamic power (2.5×). A case study for an FeFET-based LiM accumulator further demonstrates that by employing FeFET as both a switch and an NV storage element, the FeFET-based accumulator can save area (36%) and power consumption (40%) when compared with a conventional CMOS accumulator with the same structure. Xunzhao Yin, Xiaoming Chen 0003, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2018 | Computing with ferroelectric FETs: Devices, models, systems, and applicationsabstractIn this paper, we consider devices, circuits, and systems comprised of transistors with integrated ferroelectrics. Said structures are actively being considered by various semiconductor manufacturers as they can address a large and unique design space. Transistors with integrated ferroelectrics could (i) enable a better switch (i.e., offer steeper subthreshold swings), (ii) are CMOS compatible, (iii) have multiple operating modes (i.e., I-V characteristics can also enable compact, 1-transistor, non-volatile storage elements, as well as analog synaptic behavior), and (iv) have been experimentally demonstrated (i.e., with respect to all of the aforementioned operating modes). These device-level characteristics offer unique opportunities at the circuit, architectural, and system-level, and are considered here from device, circuit/architecture, and foundry-level perspectives. Ahmedullah Aziz, Evelyn T. Breyer, Xiaoming Chen 0003, Suman Datta, Sumeet Kumar Gupta, Michael Hoffmann 0008, Xiaobo Sharon Hu, Adrian M. Ionescu, Matthew Jerry, Thomas Mikolajick, Halid Mulaosmanovic, Kai Ni 0004, Michael T. Niemier, Ian O'Connor, Atanu Saha, Stefan Slesazeck, Sandeep Krishna Thirumala, Xunzhao Yin |
DATE | 19 |
| 2018 | Design and optimization of FeFET-based crossbars for binary convolution neural networksabstractBinary convolution neural networks (CNNs) have attracted much attention for embedded applications due to low hardware cost and acceptable accuracy. Nonvolatile, resistive random-access memories (RRAMs) have been adopted to build crossbar accelerators for binary CNNs. However, RRAMs still face fundamental challenges such as sneak paths, high write energy, etc. We exploit another emerging nonvolatile device-ferroelectric field-effect transistor (FeFET), to build crossbars to improve the energy efficiency for binary CNNs. Due to the three-terminal transistor structure, an FeFET can function as both a nonvolatile storage element and a controllable switch, such that both write and read power can be reduced. Simulation results demonstrate that compared with two RRAM-based crossbar structures, our FeFET-based design improves write power by 5600× and 3950×, and read power by 4.1× and 3.1×. We also tackle an important challenge in crossbar-based CNN accelerators: when a crossbar array is not large enough to hold the weights of one convolution layer, how do we partition the workload and map computations to the crossbar array? We introduce a hardware-software co-optimization solution for this problem that is universal for any crossbar accelerators. Xiaoming Chen 0003, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 2 |
| 2018 | Emerging reconfigurable nanotechnologies: can they support future electronics?abstractSeveral emerging reconfigurable technologies have been explored in recent years offering device level runtime reconfigurability. These technologies offer the freedom to choose between p- and n-type functionality from a single transistor. In order to optimally utilize the feature-sets of these technologies, circuit designs and storage elements require novel design to complement the existing and future electronic requirements. An important aspect to sustain such endeavors is to supplement the existing design flow from the device level to the circuit level. This should be backed by a thorough evaluation so as to ascertain the feasibility of such explorations. Additionally, since these technologies offer runtime reconfigurability and often encapsulate more than one functions, hardware security features like polymorphic logic gates and on-chip key storage come naturally cheap with circuits based on these reconfigurable technologies. This paper presents innovative approaches devised for circuit designs harnessing the reconfigurable features of these nanotechnologies. New circuit design paradigms based on these nano devices will be discussed to brainstorm on exciting avenues for novel computing elements. Shubham Rai, Srivatsa Rangachar Srinivasa, Patsy Cadareanu, Xunzhao Yin, Xiaobo Sharon Hu, Pierre-Emmanuel Gaillardon, Narayanan Vijaykrishnan, Akash Kumar 0001 |
ICCAD | 4 |
| 2018 | Efficient Analog Circuits for Boolean SatisfiabilityabstractEfficient solutions to nonpolynomial (NP)-complete problems would significantly benefit both science and industry. However, such problems are intractable on digital computers based on the von Neumann architecture, thus creating the need for alternative solutions to tackle such problems. Recently, a deterministic, continuous-time dynamical system (CTDS) was proposed [1] to solve a representative NP-complete problem, Boolean Satisfiability (SAT). This solver shows polynomial analog time-complexity on even the hardest benchmark k-SAT (k ≥ 3) formulas, but at an energy cost through exponentially driven auxiliary variables. This paper presents a novel analog hardware SAT solver, AC-SAT, implementing the CTDS via incorporating novel, analog circuit design ideas. AC-SAT is intended to be used as a coprocessor and is programmable for handling different problem specifications. It is especially effective for solving hard k-SAT problem instances that are challenging for algorithms running on digital machines. Furthermore, with its modular design, AC-SAT can readily be extended to solve larger size problems, while the size of the circuit grows linearly with the product of the number of variables and the number of clauses. The circuit is designed and simulated based on a 32-nm CMOS technology. Simulation Program with Integrated Circuit Emphasis (SPICE) simulation results show speedup factors of ~104on even the hardest 3-SAT problems, when compared with a state-of-the-art SAT solver on digital computers. As an example, for hard problems with N = 50 variables and M = 212 clauses, solutions are found within from a few nanoseconds to a few hundred nanoseconds. Xunzhao Yin, Behnam Sedighi, Melinda Varga, Mária Ercsey-Ravasz, Zoltán Toroczkai, Xiaobo Sharon Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2017 | Design and benchmarking of ferroelectric FET based TCAMabstractWe consider how emerging transistor technologies, specifically ferroelectric field effect transistors (or FeFETs), can realize compact and energy efficient ternary content addressable memories (TCAMs). As Moore's Law-based performance scaling trends slow, and many computational tasks of interest are now more data-centric than compute-centric, researchers are looking to improve performance/save energy by integrating efficient and compact logic/processing elements into various levels of the memory hierarchy. Potential benefits include reduced I/O traffic, energy/delay from data transfers, etc. A TCAM is an example of a logic-in-memory element that is ubiquitous in routers, caches, databases, and even neural networks. Not surprisingly, researchers continue to study how emerging technologies could lead to improved TCAMs. Recent work has considered how non-volatile (NV) memory technologies (e.g., resistive random access memory (ReRAM) or magnetic tunnel junctions (MTJs)) could best be used to construct low energy, NV TCAMs. However, acceptable Ron-Roffratios and the two terminal nature of these devices introduce energy and area overheads. Due to hysteresis in a device's I-V curve, an FeFET-based NV TCAM, offers low area overhead, as well as search energies and search speeds that are superior to other TCAM designs (i.e., based on MTJ, ReRAM and CMOS in array- and architectural-level evaluations). Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 1 |
| 2017 | An analog SAT solver based on a deterministic dynamical system: (Invited paper)abstractBoolean Satisfiability (SAT), the first problem proven to be NP-complete, is intractable on digital computers based on the von Neumann architecture. An efficient SAT solver can benefit many applications such as artificial intelligence, circuit design, and functional verification. Recently, a SAT solver approach based on a deterministic, continuous-time dynamical system (CTDS) was introduced [1]. This approach shows polynomial analog time-complexity on even the hardest k-SAT (k ≥ 3) problem instances, but at an energy cost dependent on exponentially growing auxiliary variables. This paper reports a novel analog hardware SAT solver, AC-SAT, implementing the CTDS via incorporating novel, analog circuit design ideas. AC-SAT is intended to be used as a co-processor and is programmable for handling different problem specifications. Furthermore, with its modular design, AC-SAT can be readily extended to solve larger size problems. SPICE simulation results show that AC-SAT can indeed solve the SAT problems, and it has speedup factors of ~104on even the hardest 3-SAT problems, when compared with a state-of-the-art SAT solver on digital computers. Xunzhao Yin, Zoltán Toroczkai, Xiaobo Sharon Hu |
ICCAD | 1 |
| 2016 | Using emerging technologies for hardware security beyond PUFs
Xiaobo Sharon Hu, Yier Jin, Michael T. Niemier, Xunzhao Yin |
DATE | 5 |
| 2016 | Design of latches and flip-flops using emerging tunneling devices
Xunzhao Yin, Behnam Sedighi, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 1 |
| 2016 | Enhancing Hardware Security with Emerging Transistor TechnologiesabstractWe consider how the I-V characteristics of emerging transistors (particularly those sponsored by STARnet) might be employed to enhance hardware security. An emphasis of this work is to move beyond hardware implementations of physically unclonable functions (PUFs) and random num- ber generators (RNGs). We highlight how new devices (i) may enable more sophisticated logic obfuscation for IP protection, (ii) could help to prevent fault injection attacks, (iii) prevent differential power analysis in lightweight cryptographic systems, etc. Yu Bi, Xiaobo Sharon Hu, Yier Jin, Michael T. Niemier, Kaveh Shamsi, Xunzhao Yin |
ACM Great Lakes Symposium on VLSI | 6 |
| 2016 | Exploiting ferroelectric FETs for low-power non-volatile logic-in-memory circuitsabstractNumerous research efforts are targeting new devices that could continue performance scaling trends associated with Moore's Law and/or accomplish computational tasks with less energy. One such device is the ferroelectric FET (FeFET), which offers the potential to be scaled beyond the end of the silicon roadmap as predicted by ITRS. Furthermore, the Ids vs. Vgs characteristics of FeFETs may allow a device to function as both a switch and a non-volatile storage element. We exploit this FeFET property to enable fine-grained logic-in-memory (LiM). We consider three different circuit design styles for FeFET-based LiM: complementary (differential), dynamic current mode, and dynamic logic. Our designs are compared with existing approaches for LiM (i.e., based on magnetic tunnel junctions (MTJs), CMOS, etc.) that afford the same circuit-level functionality. Assuming similar feature sizes, non-volatile FeFET-based LiM circuits are more efficient than functional equivalents based on MTJs when considering metrics such as propagation delay (2.9×, 6.8×) and dyanmic power (3.7×, 2.3×) (for 45 nm, 22 nm technology respectively). Compared to CMOS functional equivalents, FeFET designs still exhibit modest improvements in the aforementioned metrics while also offering non-volatility and reduced device count. Xunzhao Yin, Ahmedullah Aziz, Joseph Nahas, Suman Datta, Sumeet Kumar Gupta, Michael T. Niemier, Xiaobo Sharon Hu |
ICCAD | 1 |
| 2016 | Emerging Technology-Based Design of Primitives for Hardware SecurityabstractHardware security concerns such as intellectual property (IP) piracy and hardware Trojans have triggered research into circuit protection and malicious logic detection from various design perspectives. In this article, emerging technologies are investigated by leveraging their unique properties for applications in the hardware security domain. Security, for the first time, will be treated as one design metric for emerging nano-architecture. Five example circuit structures including camouflaging gates, polymorphic gates, current/voltage-based circuit protectors, and current-based XOR logic are designed to show the high efficiency of silicon nanowire FETs and graphene SymFET in applications such as circuit protection and IP piracy prevention. Simulation results indicate that highly efficient and secure circuit structures can be achieved via the use of non-CMOS devices. Yu Bi, Kaveh Shamsi, Jiann-Shiun Yuan, Pierre-Emmanuel Gaillardon, Giovanni De Micheli, Xunzhao Yin, Xiaobo Sharon Hu, Michael T. Niemier, Yier Jin |
ACM J. Emerg. Technol. Comput. Syst. | 6 |