VLDB 2026 Research / reviewers in the wild / expert
Michael T. Niemier
dblp:62/3778 · also Michael Thaddeus Niemier
· DBLP profile ↗
108ranked-venue papers
9as first author
34since 2021 · last 2026
0000-0001-7776-4306ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 101 · 9 first-author · 31 since 2021Software engineering, systems software and programming languages · 26 · 3 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 1 first-author · 5 since 2021Artificial intelligence and machine learning · 1Security and privacy · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Increasing the Efficiency of Associative Processor Architectures via CMOS-Compatible HybridizationabstractWe present a hybrid, general-purpose, associative processing-in-memory architecture that combines the energy and area advantages of a primary FeFET-based CAM array with the write performance and endurance of a much smaller CMOS-based sidekick. The hybrid nature of the architecture is transparent to the programmer, who uses a RISC-V ISA with standard RVV vector extensions. Detailed SPICE- and system-level simulations show our hybrid design dramatically curbs the endurance disadvantages of a pure FeFET design and delivers, on average, 30% and 11% area and energy savings over a purely CMOS implementation, respectively, at a performance loss of barely 1% over pure CMOS. Socrates S. Wong, Cecilio C. Tamarit, Mohammad Mehdi Sharifi, Zephan M. Enciso, Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu, José F. Martínez |
DATE | 6 |
| 2026 | Enhancing Robustness of Content-Addressable Memories for In-Memory Search
Liu Liu 0023, Tomas Sousa Pereira, Mohammad Mehdi Sharifi, Can Li 0024, Kai Ni 0004, Michael T. Niemier, Xiaobo Sharon Hu |
VTS | 8 |
| 2026 | PPIMCE: In-Memory Computing Fabric for Privacy Preserving Computing
Jianqiao Mo, Dayane Reis, Jonathan Takeshita, Taeho Jung, Brandon Reagen, Michael T. Niemier, Xiaobo Sharon Hu |
J. Comput. Sci. Technol. | 7 |
| 2026 | EvaCAM: A Circuit-Level Evaluation Tool for General Content Addressable MemoriesabstractContent addressable memories (CAMs) are special-purpose in-memory computing units that support parallel searches directly in memory. There is growing interest in CAMs for data-intensive applications such as machine learning, data mining, and bioinformatics, which has led to a rapidly growing CAM design space. CAM cells can be implemented exclusively by CMOS or with various non-volatile memory (NVM) devices. In addition to traditional binary and ternary CAMs (BCAMs and TCAMs), analog CAM (ACAM) and multi-bit CAM (MCAM) designs have recently been introduced, which could further improve density, and also support unique in-memory distance functions. Furthermore, aside from the widely-used exact match function, CAM-based approximate match functions, such as threshold match and best match, have been proposed to further extend the utility of CAMs to new application spaces. As the CAM design space is large, evaluating different CAM design options for a given application is both crucial and challenging. This paper presents EvaCAM, a circuit-level modeling and evaluation tool for CAMs. EvaCAM supports TCAM, ACAM, and MCAM designs implemented in either CMOS or NVMs, for both exact and approximate match functions. It also allows for the exploration of different CAM designs under various optimization targets. EvaCAM has been validated against measured data from fabricated chips and detailed SPICE simulations. A comprehensive design space exploration for CAMs is provided to illustrate the impact of various design decisions and to demonstrate the use cases of EvaCAM. Liu Liu 0023, Mohammad Mehdi Sharifi, Kunshi Wang, Ruibin Mao, Kai Ni 0004, Can Li 0024, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 8 |
| 2026 | Cell Instance Segmentation: The Devil Is in the BoundariesabstractState-of-the-art (SOTA) methods for cell instance segmentation are based on deep learning (DL) semantic segmentation approaches, focusing on distinguishing foreground pixels from background pixels. In order to identify cell instances from foreground pixels (e.g., pixel clustering), most methods decompose instance information into pixel-wise objectives, such as distances to foreground-background boundaries (distance maps), heat gradients with the center point as heat source (heat diffusion maps), and distances from the center point to foreground-background boundaries with fixed angles (star-shaped polygons). However, pixel-wise objectives may lose significant geometric properties of the cell instances, such as shape, curvature, and convexity, which require a collection of pixels to represent. To address this challenge, we present a novel pixel clustering method, called Ceb (for Cell boundaries), to leverage cell boundary features and labels to divide foreground pixels into cell instances. Starting with probability maps generated from semantic segmentation, Ceb first extracts potential foreground-foreground boundaries (i.e., boundary candidates) with a revised Watershed algorithm. For each boundary candidate, a boundary feature representation (called boundary signature) is constructed by sampling pixels from the current foreground-foreground boundary as well as the neighboring background-foreground boundaries. Next, a lightweight boundary classifier is used to predict its binary boundary label based on the corresponding boundary signature. Finally, cell instances are obtained by dividing or merging neighboring regions based on the predicted boundary labels. Extensive experiments on six datasets demonstrate that Ceb outperforms existing pixel clustering methods on semantic segmentation probability maps. Moreover, Ceb achieves highly competitive performance compared to state-of-the-art cell instance segmentation methods. The code is available at: https://github.com/pxliang/Ceb. Peixian Liang, Yifan Ding 0001, Yizhe Zhang 0001, Jianxu Chen 0001, Hao Zheng 0006, Yejia Zhang, Guangyu Meng, Tim Weninger, Michael T. Niemier, Xiaobo Sharon Hu, Danny Ziyi Chen |
IEEE Trans. Medical Imaging | 10 |
| 2026 | Efficient Approximation of Earth Mover's Distance Based on Nearest Neighbor Search
Guangyu Meng, Ruyu Zhou, Liu Liu 0023, Peixian Liang, Fang Liu 0006, Danny Ziyi Chen, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Multim. | 7 |
| 2025 | Towards Uncertainty-aware Robotic Perception via Mixed-signal BNN Engine Leveraging Probabilistic Quantum TunnelingabstractIntegrating deep learning with environmental perception enhances robotic adaptability to complex tasks. However, its “black-box” nature, such as the lack of uncertainty quantification, poses challenges for safety-critical applications, particularly in unstructured and noisy environments. Bayesian neural networks (BNNs) offer uncertainty quantification but are limited by high hardware overhead, restricting real-time implementation on resource-constrained robots. This paper presents a mixedsignal hardware accelerator for BNNs, utilizing probabilistic quantum tunneling in fully depleted silicon-on-insulator (FDSOI) transistors to enable efficient, real-time uncertainty quantification. Device measurements indicate high-quality Gaussian random variable generation, validated through quantile-quantile plot analysis, with a high correlation coefficient ($r=0.997$) at $200 \mathrm{fJ} /$ sample. Leveraging such compact randomness, the parallel architecture achieved $10^{3}-10^{4} \times$ latency reduction at less than $2 \times$ area cost. Finally, in uncertainty-aware visual localization application of autonomous underwater vehicles, the BNN model effectively distinguishes data noise from model uncertainty, yielding significant information gain and enhancing the resampling efficiency by $4.5 \times$ at same accuracy. Likai Pei, Xingtian Wang, Xueji Zhao, Wanxin Huang, Boyang Cheng, Halid Mulaosmanovic, Stefan Dünkel, Dominik Kleimaier, Sven Beyer, Kai Ni 0004, Mengxue Hou, Michael T. Niemier, Ningyuan Cao |
DAC | 13 |
| 2025 | COSMOS: RL-Enhanced Locality-Aware Counter Cache Optimization for Secure MemoryabstractSecure memory systems employing AES-CTR encryption face significant performance challenges due to high counter (CTR) cache miss rates, especially in applications with irregular memory access patterns.These high miss rates increase memory traffic and latency, as each CTR cache miss triggers additional DRAM accesses.To address these bottlenecks and adapt to diverse access patterns, we propose COSMOS (Counter Optimized Secure Memory Operation Scheme), a novel solution leveraging reinforcement learning to reduce long memory access latency.COSMOS integrates two RL-based specialized predictors: one for data location prediction and another for CTR locality prediction, each with a well-defined state space, action space, and reward function.The RL-based data location predictor determines whether data reside on-chip or offchip after an L1 cache miss, enabling early CTR access for off-chip predictions with minimal changes to the existing cache hierarchy.The RL-based CTR locality predictor identifies CTRs with high locality, supporting a locality-centric CTR cache (LCR-CTR) to improve cache efficiency and reduce miss rates.COSMOS improves performance over MorphCtr by 25% in for irregular memory access applications, with minimal hardware overhead. Xiaoyang Lu, Yuezhi Che, Ziang Tian, Dazhao Cheng, Xian-He Sun, Michael T. Niemier, Xiaobo Sharon Hu |
MICRO | 7 |
| 2025 | Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAMabstractProcessing-in-Memory (PIM) enhances memory with computational capabilities, potentially solving energy and latency issues associated with data transfer between memory and processors. However, managing concurrent computation and data flow within the PIM architecture incurs significant latency and energy penalty for applications. This paper introduces Shared-PIM, an architecture for in-DRAM PIM that strategically allocates rows in memory banks, bolstered by memory peripherals, for concurrent processing and data movement. Shared-PIM enables simultaneous computation and data transfer within a memory bank. When compared to LISA, a state-of-the-art architecture that facilitates data transfers for in-DRAM PIM, Shared-PIM reduces data movement latency and energy by 5× and 1.2×, respectively. Furthermore, when integrated to a state-of-the-art (SOTA) in-DRAM PIM architecture (pLUTo), Shared-PIM achieves 1.4× faster addition and multiplication, and thereby improves the performance of matrix multiplication (MM) tasks by 40%, polynomial multiplication (PMM) by 44%, and numeric number transfer (NTT) tasks by 31%. Moreover, for graph processing tasks like Breadth-First Search (BFS) and Depth-First Search (DFS), Shared-PIM achieves a 29% improvement in speed, all with an area overhead of just 7.16% compared to the baseline pLUTo. Ahmed Mamdouh, Michael T. Niemier, Xiaobo Sharon Hu, Dayane Reis |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 3 |
| 2024 | Smoothing Disruption Across the Stack: Tales of Memory, Heterogeneity, & CompilersabstractMultiple research vectors represent possible paths to improved energy and performance metrics at the application-level. There are active efforts with respect to emerging logic devices, new memory technologies, novel interconnects, and heterogeneous integration architectures. Of great interest is quantifying the potential impact of a given solution to prioritize research vectors accordingly. In this paper, we discuss two efforts - one focused on emerging memory technology, and another focused on heterogeneous integration technology - that speak to best practices for, and needed contributions from the design automation (DA) community to explore this vast design space. Furthermore, we highlight new research efforts that aim to develop the novel compiler abstractions and frameworks that are ultimately needed to derive maximum value from new memory and/or heterogeneous and monolithic integration architecture, and that can also play an important role with respect to design space exploration efforts. Michael T. Niemier, Zephan M. Enciso, M. Sharifi, Xiaobo Sharon Hu, Ian O'Connor, A. Graening, Jerónimo Castrillón, João Paulo C. de Lima, Asif Ali Khan, Hamid Farzaneh, N. Afroze, Julien Ryckaert |
DATE | 1 |
| 2024 | TAP-CAM: A Tunable Approximate Matching Engine based on Ferroelectric Content Addressable MemoryabstractPattern search is crucial in numerous analytic applications for retrieving data entries akin to the query. Content Addressable Memories (CAMs), an in-memory computing fabric, directly compare input queries with stored entries through embedded comparison logic, facilitating fast parallel pattern search in memory. While conventional CAM designs offer exact match functionality, they are inadequate for meeting the approximate search needs of emerging data-intensive applications. Some recent CAM designs propose approximate matching functions, but they face limitations such as excessively large cell area or the inability to precisely control the degree of approximation. In this paper, we propose TAP-CAM, a novel ferroelectric field effect transistor (FeFET) based ternary CAM (TCAM) capable of both exact and tunable approximate matching. TAP-CAM employs a compact 2FeFET-2R cell structure as the entry storage unit, and similarities in Hamming distances between input queries and stored entries are measured using an evaluation transistor associated with the matchline of CAM array. The operation, robustness and performance of the proposed design at array level have been discussed and evaluated, respectively. We conduct a case study of K-nearest neighbor (KNN) search to benchmark the proposed TAP-CAM at application level. Results demonstrate that compared to 16T CMOS CAM with exact match functionality, TAP-CAM achieves a 16.95× energy improvement, along with a 3.06% accuracy enhancement. Compared to 2FeFET TCAM with approximate match functionality, TAP-CAM achieves a 6.78× energy improvement. Chenyu Ni, Che-Kai Liu, Liu Liu 0023, Mohsen Imani, Thomas Kämpfe, Kai Ni 0004, Michael T. Niemier, Xiaobo Sharon Hu, Cheng Zhuo, Xunzhao Yin |
ICCAD | 8 |
| 2024 | Towards Uncertainty-Quantifiable Biomedical Intelligence: Mixed-signal Compute-in-Entropy for Bayesian Neural NetworksabstractTo enhance AI robustness of mission-critical biomedical applications, Bayesian Neural Networks (BNNs) are instrumental for their structured approach to AI uncertainty estimation. However, implementing BNNs on edge devices is challenging due to significant resource demands for dynamic model updates and extensive inference sampling. Addressing this, we introduce a novel mixed-signal Compute-in-Memory with Entropy (CIE) hardware architecture that segregates dynamically-generated weights into analog entropy and digital parameters within a compute-in-memory framework, greatly reducing hardware overhead. We conducted thorough evaluations of the CIE architecture, assessing its performance against varying hardware imperfections, such as digital quantization errors, analog distribution imperfections, and device process variations, with a focus on both general and specialized tasks like Ventricular Arrhythmia (VA) detection. Our contributions include (1) a generic BNN acceleration strategy suitable for various CIM techniques and emerging devices, (2) a custom circuit design that improves hardware efficiency by 19.2×-440× compared to existing BNN accelerators, (3) a CIE-based BNN for VA detection enhancing accuracy, reducing uncertainty estimation time and energy/latency to 1.29μJ/1.55ms, and (4) identification of tolerable quantization error and device variation limits for BNNs in uncertainty estimation. Likai Pei, Zephan M. Enciso, Boyang Cheng, Steven Davis, Zhenge Jia, Michael T. Niemier, Yiyu Shi 0001, Xiaobo Sharon Hu, Ningyuan Cao |
ICCAD | 8 |
| 2024 | Design of High-Performance and Compact CAM for Supporting Data-Intensive ApplicationsabstractContent addressable memory (CAM) is a special-purpose search engine that can support parallel search directly in memory. CAMs are of increasing interest for machine learning and data analytics applications that require intensive search operations. However, conventional CMOS CAMs have large cell areas and high energy consumption, which limits applicability. Also, many data-intensive applications need more efficient data representation and approximate matching functions, which may not be efficiently realized by conventional ternary CAMs. As such, we introduce a more compact and high-performance CAM design based on non-volatile ferroelectic FET devices. Furthermore, we present a reconfigurable CAM design, MHCAM, to support approximate search for multi-dimensional data. We use DNA alignment as a proxy application to illustrate the design’s application-level benefits. Liu Liu 0023, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu |
ISCAS | 3 |
| 2024 | Accelerating Finite-Field and Torus Fully Homomorphic Encryption via Compute-Enabled (S)RAMabstractFully Homomorphic Encryption (FHE) allows outsourced computation on clients’ encrypted data while preserving data privacy. FHE’s high computational intensity incurs high overhead from data transfer with hardware such as CPU, GPU, and FPGA, due to the inherent separation between computing and data. To overcome this limitation, Compute-Enabled RAM (CE-RAM) has been explored; however, prior work using CE-RAM to accelerate FHE only explores a simple implementation of a finite-field FHE scheme and did not explore algorithmic optimizations.In this paper, we investigate CE-RAM acceleration FHE more deeply, implementing both the finite-field B/FV and torus-based TFHE cryptosystems in CE-RAM with common FHE optimizations. This is the first work to explore using CE-RAM to accelerate TFHE. For B/FV, we explore parameter-specific algorithmic optimizations specifically designed for CE-RAM friendliness. We evaluate our implementation as compared to prior work in CE-RAM FHE acceleration and other hardware acceleration strategies. We demonstrate speedups of up to 784x for B/FV homomorphic multiplication and 38x for TFHE bootstrapping as compared to CPU implementations. We also discuss the overhead of CE-RAM for FHE on energy and area consumption, showing comparable or improved performance as compared to other work or hypothetical near-memory accelerators. Jonathan Takeshita, Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu, Taeho Jung |
IEEE Trans. Computers | 4 |
| 2023 | In-Memory Computing Accelerators for Emerging Learning ParadigmsabstractOver the past decades, emerging, data-driven machine learning (ML) paradigms have increased in popularity, and revolutionized many application domains. To date, a substantial effort has been devoted to devising mechanisms for facilitating the deployment and near ubiquitous use of these memory intensive ML models. This review paper presents the use of in-memory computing (IMC) accelerators for emerging ML paradigms from a bottom-up perspective through the choice of devices, the design of circuits/architectures, to the application-level results. Dayane Reis, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu |
ASP-DAC | 3 |
| 2023 | Cross Layer Design for the Predictive Assessment of Technology-Enabled ArchitecturesabstractThere is great interest in “end-to-end” analysis that captures how innovation at the materials, device, and/or archi-tectural levels will impact figures of merit at the application-level. However, there are numerous combinations of devices and architectures to study, and we must establish systematic ways to accurately explore and cull a vast design space. We aim to capture how innovations at the materials/device-level may ultimately impact figures of merit associated with both existing and emerging technologies that may be employed for either logic and/or memory. We will highlight how collaborations with researchers at these levels of the design hierarchy - as well as efforts to help construct well-calibrated device models - can in-turn support architectural design space explorations that will help to identify the most promising ways to use new technologies to support application-level workloads of interest. For given compute workloads, we can then quantitatively assess the potential benefits of technology-driven architectures to identify the most promising paths forward. Because of the large number of potentially interesting device-architecture combinations, it is of the utmost importance to develop well-calibrated analytical modeling tools to more rapidly assess the potential value of a given (likely heterogeneous) solution. We highlight recent efforts and needs in this space. Michael T. Niemier, Xiaobo Sharon Hu, Liu Liu 0023, Mohammad Mehdi Sharifi, Ian O'Connor, David Atienza 0001, Giovanni Ansaloni, Can Li 0024, Daniel C. Ralph |
DATE | 1 |
| 2023 | Invited Paper: Algorithm/Hardware Co-Design for Few-Shot Learning at the EdgeabstractOn-device learning is essential to achieve intelligence at the edge, where it is desirable to learn from few samples or even just a single sample. Memory-augmented neural networks (MANNs), which augment neural networks with an attentional memory, can draw on already learnt knowledge patterns and adapt to new but similar tasks. Implementing MANNs on conventional architectures can require a significant amount of costly data transfer, thereby limiting the practical use of MANNs at the edge. In this paper, we introduce algorithm/hardware co-design solutions which exploit compact designs of content addressable memories (CAMs) based on emerging non-volatile memories (e.g., FeFETs) to implement energy-efficient MANN accelerators. The design space of MANN accelerators is systematically analyzed by considering different circuit, architecture, and algorithm options. We further discuss how hyper-dimensional representations of data can be combined with MANNs to overcome the negative effect of device/circuit variabilities on learning quality, thus achieving not only energy-efficient but also accuracy-competitive on-device learning at the edge. We also investigate modeling of device-to-device (D2D) variation in FeFETs using the write-with-verify approach and detail its impact on the energy, delay, and accuracy of the MANN application. Ann Franchesca Laguna, Mohammad Mehdi Sharifi, Dayane Reis, Liu Liu 0023, Andrew Hennessee, Clayton O'Dell, Ian O'Connor, Michael T. Niemier, Xiaobo Sharon Hu |
ICCAD | 8 |
| 2023 | Accelerating Polynomial Modular Multiplication with Crossbar-Based Compute-in-MemoryabstractLattice-based cryptographic algorithms built on ring learning with error theory are gaining importance due to their potential for providing post-quantum security. However, these algorithms involve complex polynomial operations, such as polynomial modular multiplication (PMM), which is the most time-consuming part of these algorithms. Accelerating PMM is crucial to make lattice-based cryptographic algorithms widely adopted by more applications. This work introduces a novel high-throughput and compact PMM accelerator, X-Poly, based on the crossbar (XB)-type compute-in-memory (CIM). We identify the most appropriate PMM algorithm for XB-CIM. We then propose a novel bit-mapping technique to reduce the area and energy of the XB-CIM fabric, and conduct processing engine (PE)-level optimization to increase memory utilization and support different problem sizes with a fixed number of XB arrays. X-Poly design achieves$3.1\times 10^{6}$PMM operations/s throughput and offers 200 x latency improvement compared to the CPU-based implementation. It also achieves 3.9 x throughput per area improvement compared with the state-of-the-art CIM accelerators. Mengyuan Li 0001, Michael T. Niemier, Xiaobo Sharon Hu |
ICCAD | 3 |
| 2023 | A Reconfigurable FeFET Content Addressable Memory for Multi-State Hamming DistanceabstractPattern searches, a key operation in many data analytic applications, often deal with data represented by multiple states per dimension. However, hash tables, a common software-based pattern search approach, require a large amount of additional memory, and thus, are limited by the memory wall. A hardware-based solution is to use content-addressable memories (CAMs) that support fast associative searches in parallel. Ternary CAMs (TCAMs) support bit-wise Hamming distance (HD) based searches. Detecting the HD of vectors with multiple states per dimension (i.e., multi-state Hamming distance (MSHD)) can be implemented on TCAMs with one-hot encoding, but requires one TCAM cell per state, leading to a higher area, latency, and energy overhead. We propose a Ferroelectric FET (FeFET)-based multi-state CAM design, MHCAM, which implements MSHD searches in a dense FeFET-based memory array. MHCAM only uses$\lceil log_{2} s \rceil ~2$FeFET CAM cells to represent$s$states or symbols per dimension, and can be reconfigured to 2-bit/4-bit/6-bit/8-bit dimensions. A low-cost sensing circuit with matchline voltage scaling technique is introduced to perform both exact match and threshold match. We use DNA and protein pre-alignment filtering as application case studies to evaluate the application-level benefit of MHCAM. DNA and protein pre-alignment filtering achieve$3.8\times /4.7\times $speedup and$1.7\times /1.8\times $energy improvement compared with the state-of-the-art 2FeFET TCAM-based implementation. Liu Liu 0023, Ann Franchesca Laguna, Ramin Rajaei, Mohammad Mehdi Sharifi, Arman Kazemi, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Circuits Syst. I Regul. Pap. | 7 |
| 2022 | iMARS: an in-memory-computing architecture for recommendation systemsabstractRecommendation systems (RecSys) suggest items to users by predicting their preferences based on historical data. Typical RecSys handle large embedding tables and many embedding table related operations. The memory size and bandwidth of the conventional computer architecture restrict the performance of RecSys. This work proposes an in-memory-computing (IMC) architecture (iMARS) for accelerating the filtering and ranking stages of deep neural network-based RecSys. iMARS leverages IMC-friendly embedding tables implemented inside a ferroelectric FET based IMC fabric. Circuit-level and system-level evaluation show that iMARS achieves 16.8x (713x) end-to-end latency (energy) improvement compared to the GPU counterpart for the MovieLens dataset. Mengyuan Li 0001, Ann Franchesca Laguna, Dayane Reis, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
DAC | 5 |
| 2022 | Eva-CAM: A Circuit/Architecture-Level Evaluation Tool for General Content Addressable MemoriesabstractContent addressable memories (CAMs), a special-purpose in-memory computing (IMC) unit, support parallel searches directly in memory. There are growing interests in CAMs for data-intensive applications such as machine learning and bioinformatics. The design space for CAMs is rapidly expanding. In addition to traditional ternary CAMs (TCAMs), analog CAM (ACAM) and multi-bit CAM (MCAM) designs based on various non-volatile memory (NVM) devices have been recently introduced and may offer higher density, better energy efficiency, and non-volatility. Furthermore, aside from the widely-used exact match based search, CAM-based approximate matches have been proposed to further extend the utility of CAMs to new application spaces. For this memory architecture, evaluating different CAM design options for a given application is becoming more challenging. This paper presents Eva-CAM, a circuit/architecture-level modeling and evaluation tool for CAMs. Eva-CAM supports TCAM, ACAM, and MCAM designs implemented in non-volatile memories, for both exact and approximate match types. It also allows for the exploration of CAM array structures and sensing circuits. Eva-CAM has been validated with HSPICE simulation results and chip measurements. A comprehensive case study is described for FeFET CAM design space exploration. Liu Liu 0023, Mohammad Mehdi Sharifi, Ramin Rajaei, Arman Kazemi, Kai Ni 0004, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 7 |
| 2022 | COSIME: FeFET Based Associative Memory for In-Memory Cosine Similarity SearchabstractIn a number of machine learning models, an input query is searched across the trained class vectors to find the closest feature class vector in cosine similarity metric. However, performing the cosine similarities between the vectors in Von-Neumann machines involves a large number of multiplications, Euclidean normalizations and division operations, thus incurring heavy hardware energy and latency overheads. Moreover, due to the memory wall problem that presents in the conventional architecture, frequent cosine similarity-based searches (CSSs) over the class vectors requires a lot of data movements, limiting the throughput and efficiency of the system. To overcome the aforementioned challenges, this paper introduces COSIME, a general in-memory associative memory (AM) engine based on the ferroelectric FET (FeFET) device for efficient CSS. By leveraging the one-transistor AND gate function of FeFET devices, current-based translinear analog circuit and winner-take-all (WTA) circuitry, COSIME can realize parallel in-memory CSS across all the entries in a memory block, and output the closest word to the input query in cosine similarity metric. Evaluation results at the array level suggest that the proposed COSIME design achieves 333× and 90.5× latency and energy improvements, respectively, and realizes better classification accuracy when compared with an AM design implementing approximated CSS. The proposed in-memory computing fabric is evaluated for an HDC problem, showcasing that COSIME can achieve on average 47.1× and 98.5× speedup and energy efficiency improvements compared with an GPU implementation. Che-Kai Liu, Haobang Chen, Mohsen Imani, Kai Ni 0004, Arman Kazemi, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu, Liang Zhao 0004, Cheng Zhuo, Xunzhao Yin |
ICCAD | 7 |
| 2022 | FeFET Multi-Bit Content-Addressable Memories for In-Memory Nearest Neighbor SearchabstractNearest neighbor (NN) search computations are at the core of many applications such as few-shot learning, classification, and hyperdimensional computing. As such, efficient hardware support for NN search is highly desired. In-memory computing using emerging devices offers attractive solutions for NN search. Solutions based on ternary content-addressable memories (TCAMs) offer high energy and latency improvements for NN search at the expense of accuracy. In this work, we propose a novel distance function that can be natively evaluated with multi-bit content-addressable memories (MCAMs) based on ferroelectric FETs (FeFETs) to perform a single-step, in-memory NN search. We evaluate the efficacy of FeFET MCAMs in the context of few-shot learning applications with different datasets. As an example, we achieve a 78.54% accuracy for a 5-way, 5-shot classification task for the mini-ImageNet dataset (only 1.5% lower than software-based implementations) when using a 3-bit MCAM for NN search. We consider the effects of FeFET threshold voltage variations on the application accuracy and analyze the area and search energy requirements of FeFET MCAMs for accurate operations. Our results indicate that MCAMs require 2× lower area and search energy than TCAMs to achieve the same accuracy. Furthermore, we experimentally demonstrate a 2-bit implementation of FeFET MCAM using AND arrays from GLOBALFOUNDRIES to further validate the design concept. Arman Kazemi, Mohammad Mehdi Sharifi, Ann Franchesca Laguna, Franz Müller 0001, Xunzhao Yin, Thomas Kämpfe, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Computers | 7 |
| 2022 | IMCRYPTO: An In-Memory Computing Fabric for AES Encryption and DecryptionabstractThis article proposes IMCRYPTO, an in-memory computing (IMC) fabric for accelerating advanced encryption standard (AES) encryption and decryption. IMCRYPTO employs a unified structure to implement encryption and decryption in a single-hardware architecture with combined (Inv)SubBytes and (Inv)MixColumns steps. Because of this step combination and the high parallelism achieved by multiple units of random access memory (RAM) and random access/content addressable memory (RA/CAM) arrays, IMCRYPTO achieves high-throughput encryption and decryption without sacrificing area and power consumption. In addition, due to the integration of an RISC-V core, IMCRYPTO offers programmability and flexibility. IMCRYPTO improves the throughput per area by a minimum (maximum) of$3.3\times $($223.1\times $) compared to previous ASICs/IMC architectures for AES-128 encryption. Projections show added benefit from emerging technologies of up to$5.3\times $to the area–delay–power product of IMCRYPTO. Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2021 | Cross-layer Design for Computing-in-Memory: From Devices, Circuits, to Architectures and ApplicationsabstractThe era of Big Data, Artificial Intelligence (AI) and Internet of Things (IoT) is approaching, but our underlying computing infrastructures are not sufficiently ready. The end of Moore's law and process scaling as well as the memory wall associated with von Neumann architectures have throttled the rapid development of conventional architectures based on CMOS technology, and cross-layer efforts that involve the interactions from low-end devices to high-end applications have been prominently studied to overcome the aforementioned challenges. On one hand, various emerging devices, e.g., Ferroelectric FET, have been proposed to either sustain the scaling trends or enable novel circuit and architecture innovations. On the other hand, novel computing architectures/algorithms, e.g., computing-in-memory (CiM), have been proposed to address the challenges faced by conventional von Neumann architectures. Naturally, integrated approaches across the emerging devices and computing architectures/algorithms for data-intensive applications are of great interests. This paper uses the FeFET as a representative device, and discuss about the challenges, opportunities and contributions for the emerging trends of cross-layer co-design for CiM. Hussam Amrouch, Xiaobo Sharon Hu, Mohsen Imani, Ann Franchesca Laguna, Michael T. Niemier, Simon Thomann, Xunzhao Yin, Cheng Zhuo |
ASP-DAC | 5 |
| 2021 | Attention-in-Memory for Few-Shot Learning with Configurable Ferroelectric FET ArraysabstractAttention-in-Memory (AiM), a computing-in-memory (CiM) design, is introduced to implement the attentional layer of Memory Augmented Neural Networks (MANNs). AiM consists of a memory array based on Ferroelectric FETs (FeFET) along with CMOS peripheral circuits implementing configurable functionalities, i.e., it can be dynamically changed from a ternary content-addressable memory (TCAM) to a general-purpose (GP) CiM. When compared to state-of-the art accelerators, AiM achieves comparable end-to-end speed-up and energy for MANNs, with better accuracy (95.14% v.s. 92.21%, and 95.14% v.s. 91.98%) at iso-memory size, for a 5-way 5-shot inference task with the Omniglot dataset. Dayane Reis, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu |
ASP-DAC | 3 |
| 2021 | In-Memory Nearest Neighbor Search with FeFET Multi-Bit Content-Addressable MemoriesabstractNearest neighbor (NN) search is an essential operation in many applications, such as one/few-shot learning and image classification. As such, fast and low-energy hardware support for accurate NN search is highly desirable. Ternary content-addressable memories (TCAMs) have been proposed to accelerate NN search for few-shot learning tasks by implementing$L$∞and Hamming distance metrics, but they cannot achieve software-comparable accuracies. This paper proposes a novel distance function that can be natively evaluated with multi-bit content-addressable memories (MCAMs) based on ferroelectric FETs (Fe-FETs) to perform a single-step, in-memory NN search. Moreover, this approach achieves accuracies comparable to floating-point precision implementations in software for NN classification and one/few-shot learning tasks. As an example, the proposed method achieves a 98.34% accuracy for a 5-way, 5-shot classification task for the Omniglot dataset (only 0.8% lower than software-based implementations) with a 3-bit MCAM. This represents a 13% accuracy improvement over state-of-the-art TCAM-based implementations at iso-energy and iso-delay. The presented distance function is resilient to the effects of FeFET device-to-device variations. Furthermore, this work experimentally demonstrates a 2-bit implementation of FeFET MCAM using AND arrays from GLOBALFOUNDRIES to further validate proof of concept. Arman Kazemi, Mohammad Mehdi Sharifi, Ann Franchesca Laguna, Franz Müller 0001, Ramin Rajaei, Ricardo Olivo, Thomas Kämpfe, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 8 |
| 2021 | In-Memory Computing based Accelerator for Transformer Networks for Long SequencesabstractTransformer networks have outperformed recurrent neural networks and convolutional neural networks in various sequential tasks. However, scaling transformer networks for long sequences has been challenging because of memory and compute bottlenecks. Transformer networks are impeded by memory bandwidth limitations because of their low operation per byte ratio resulting in low utilization of GPU's computing resources. In-memory processing can mitigate memory bottlenecks by eliminating the transfer time between memory and compute units. Furthermore, transformer networks use neural attention mechanisms to characterize the relationships between sequence elements. Efficient hardware solutions have been proposed to implement efficient attention mechanisms, which include ternary content addressable memories (TCAM), crossbar arrays (XBars), and processing in-memory (PIM). However, these solutions do not implement a multi-head self-attention mechanism. We propose using a combination of XBars and CAMs to accelerate transformer networks. We improve the speed of transformer networks by (1) computing in-memory, thus minimizing the memory transfer overhead, (2) caching reusable parameters to reduce the number of operations, (3) exploiting the available parallelism in the attention mechanism, and (4) using locality sensitive hashing to filter the number of sequence elements by their importance. Our approach achieves a 200x speedup and 41x energy improvement for a sequence length of 4098. Ann Franchesca Laguna, Arman Kazemi, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 3 |
| 2021 | Exploiting FeFETs via Cross-Layer Design from In-memory Computing Circuits to Meta-Learning ApplicationsabstractA ferroelectric FET (FeFET), made by integrating a ferroelectric material layer in the gate stack of a MOSFET, is a device that can behave as both a transistor and a non-volatile storage element. This unique property of FeFETs enables area efficient and low-power merged logic and memory functionality, desirable for many data analytic and machine learning applications. To best exploit this unique feature of FeFETs, cross-layer design practices spanning from circuits and architectures to algorithms and applications is needed. The paper presents FeFET-based circuits and architectures that offer, either independently or in a configurable fashion, content addressable memory (TCAM) and general-purpose compute-in-memory (GP-CiM) functionalities. These in-memory computing modules bring new opportunities to accelerating data-intensive applications. We discuss the use of these FeFET based in-memory computing fabrics in meta-learning applications, specifically as attentional memory. System-level task mapping and end-to-end evaluation will be discussed. Dayane Reis, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 3 |
| 2021 | ICCAD Tutorial Session Paper Ferroelectric FET Technology and Applications: From Devices to SystemsabstractThe rapidly increasing volume and complexity of data is demanding the relentless scaling of computing power. With transistor feature size approaching physical limits, the benefits that CMOS technology can provide is diminishing. For future energy efficient computing systems, researchers aim to exploit various emerging nanotechnologies to replace conventional CMOS technology. In particular, ferroelectric FETs (FeFETs) appear to be a promising candidate to continue improving energy efficiency for data-intensive applications. Advances in FeFET scalability and FeFET compatibility with CMOS have sparked growing interest in device, circuit, and system communities. While FeFET is still evolving, many researchers and developers are already cautiously optimistic about its future. This paper provides a review on FeFET's recent technology advances, challenges, and opportunities, with a particular emphasis upon device modeling and circuit design of FeFET content addressable memory, as well as their applications in machine learning. Hussam Amrouch, Xiaobo Sharon Hu, Arman Kazemi, Ann Franchesca Laguna, Kai Ni 0004, Michael T. Niemier, Mohammad Mehdi Sharifi, Simon Thomann, Xunzhao Yin, Cheng Zhuo |
ICCAD | 7 |
| 2021 | Low-Cost Sequential Logic Circuit Design Considering Single Event Double-Node Upsets and Single Event TransientsabstractAs CMOS device sizes continue to scale down, radiation-related reliability issues are of ever-growing concern. Single event double node upsets (SEDUs) in sequential logic and single event transients (SETs) in combinational logic are sources of high rate radiation-induced soft errors that can affect the functionality of logic circuits. This paper presents effective circuit-level solutions for combating SEDUs/SETs in nanoscale sequential and combinational logic circuits. More specifically, we propose and evaluate low-power latch and flip-flop circuits to mitigate SEDUs and SETs. Simulations with a 22 nm PTM model reveal that the proposed circuits offer full immunity against SEDUs, can better filter SET pulses, and simultaneously reduce design overhead when compared to prior work. As a representative example, simulation-based studies show that our designs offer up to 77% improvements in delay-power-area product, and can filter out up to 58% wider SET pulses when compared to the state-of-the-art. Ramin Rajaei, Michael T. Niemier, Xiaobo Sharon Hu |
ICCD | 2 |
| 2021 | A Flash-Based Multi-Bit Content-Addressable Memory with Euclidean Squared DistanceabstractContent-addressable memories (CAMs) can perform fast and energy-efficient search operations. Recently, ternary CAMs (TCAMs) have been utilized to measure Hamming distance for machine learning applications, where they offer significant energy savings and speed-ups. However, the binary precision of the Hamming distance can lead to severe degradation in application-level accuracies, thus mitigating the impact of gains with respect to other figures of merit. To enhance accuracy, multi-bit CAMs (MCAMs) have been proposed that offer higher density and energy savings than TCAMs by storing multiple bits in each cell. However, existing MCAMs are based on emerging nonvolatile memory technologies that are yet to be established. To this end, we propose a fast and extremely energy-efficient MCAM based on mature and widely used flash cells, called $\mathrm{E}^{2} -$MCAM. $\mathrm{E}^{2} -$MCAM can measure the Euclidean squared distance between search queries and data stored in the MCAM “in-memory”, and in a single cycle. We evaluate $\mathrm{E}^{2} -$MCAM using an experimentally calibrated flash model in HSPICE with 3-bit precision for proof of concept demonstration. $\mathrm{A}64 \times 32 \mathrm{E}^{2} -$MCAM array achieves a 0.34 fJ energy per bit per search and a 2.7 ns latency while operating at a $770 \mu \mathrm{W}$ power. Fast and efficient hardware support for Euclidean squared distance is highly valuable as it is widely used in a plethora of machine learning applications. As an example, we show that $\mathrm{E}^{2} -$MCAM achieves accuracies comparable to floating-point GPU implementations with only 3-bit precision for few-shot learning tasks with the ImageNet dataset while offering improvements in energy and latency. Arman Kazemi, Shubham Sahay, Ayush Saxena, Mohammad Mehdi Sharifi, Michael T. Niemier, Xiaobo Sharon Hu |
ISLPED | 5 |
| 2021 | MIMHD: Accurate and Efficient Hyperdimensional Inference Using Multi-Bit In-Memory ComputingabstractHyperdimensional Computing (HDC) is an emerging computational framework that mimics important brain functions by operating over high-dimensional vectors, called hypervectors (HVs). In-memory computing implementations of HDC are desirable since they can significantly reduce data transfer overheads. All existing in-memory HDC platforms consider binary HVs where each dimension is represented with a single bit. However, utilizing multi-bit HVs allows HDC to achieve acceptable accuracies in lower dimensions which in turn leads to higher energy efficiencies. Thus, we propose a highly accurate and efficient multi-bit in-memory HDC inference platform called MIMHD. MIMHD supports multi-bit operations using ferroelectric field-effect transistor (FeFET) crossbar arrays for multiply-and-add and FeFET multi-bit content-addressable memories for associative search. We also introduce a novel hardware-aware retraining framework (HWART) that trains the HDC model to learn to work with MIMHD. For six popular datasets and 4000 dimension HVs, MIMHD using 3-bit (2-bit) precision HVs achieves (i) average accuracies of 92.6% (88.9%) which is 8.5% (4.8%) higher than binary implementations; (ii) 84.1× (78.6×) energy improvement over a GPU, and (iii) 38.4×(34.3×) speedup over a GPU, respectively. The 3-bit MIMHD is 4.3× and 13× faster and more energy-efficient than binary HDC accelerators while achieving similar accuracies. Arman Kazemi, Mohammad Mehdi Sharifi, Zhuowen Zou, Michael T. Niemier, Xiaobo Sharon Hu, Mohsen Imani |
ISLPED | 4 |
| 2021 | Application-driven Design Exploration for Dense Ferroelectric Embedded Non-volatile MemoriesabstractThe memory wall bottleneck is a key challenge across many data-intensive applications. Multi-level FeFET-based embedded non-volatile memories are a promising solution for denser and more energy-efficient on-chip memory. However, reliable multi-level cell storage requires careful optimizations to minimize the design overhead costs. In this work, we investigate the interplay between FeFET device characteristics, programming schemes, and memory array architecture, and explore different design choices to optimize performance, energy, area, and accuracy metrics for critical data-intensive workloads. From our cross-stack design exploration, we find that we can store DNN weights and social network graphs at a density of over 8MB/mm2and sub-2ns read access latency without loss in application accuracy. Mohammad Mehdi Sharifi, Lillian Pentecost, Ramin Rajaei, Arman Kazemi, Qiuwen Lou, Gu-Yeon Wei, David Brooks 0001, Kai Ni 0004, Xiaobo Sharon Hu, Michael T. Niemier, Marco Donato |
ISLPED | 10 |
| 2020 | A Device Non-Ideality Resilient Approach for Mapping Neural Networks to Crossbar ArraysabstractWe propose a technology-independent method, referred to as adjacent connection matrix (ACM), to efficiently map signed weight matrices to non-negative crossbar arrays. When compared to same-hardware-overhead mapping methods, using ACM leads to improvements of up to 20% in training accuracy for ResNet-20 with the CIFAR-10 dataset when training with 5-bit precision crossbar arrays or lower. When compared with strategies that use two elements to represent a weight, ACM achieves comparable training accuracies, while also offering area and read energy reductions of 2.3× and 7×, respectively. ACM also has a mild regularization effect that improves inference accuracy in crossbar arrays without any retraining or costly device/variation-aware training. Arman Kazemi, Cristobal Alessandri, Alan C. Seabaugh, Xiaobo Sharon Hu, Michael T. Niemier, Siddharth Joshi 0001 |
DAC | 5 |
| 2020 | Emerging Neural Workloads and Their Impact on HardwareabstractWe consider existing and emerging neural workloads, and what hardware accelerators might be best suited for said workloads. We begin with a discussion of analog crossbar arrays, which are known to be well-suited for matrix-vector multiplication operations that are commonplace in existing neural network models such as convolutional neural networks (CNNs). We highlight candidate crosspoint devices, what device and materials challenges must be overcome for a given device to be employed in a crossbar array for a computationally interesting neural workload, and how circuit and algorithmic optimizations may be employed to mitigate undesirable characteristics from devices/materials. We then discuss two emerging neural workloads. We first consider machine learning models for one- and few-shot learning tasks (i.e., where a network can be trained with just one or a few, representative examples of a given class). Notably crossbar-based architectures can be used to accelerate said models. Hardware solutions based on content addressable memory arrays will also be discussed. We then consider machine learning models for recommendation systems. Recommendation models, an emerging class of machine learning models, employ distinct neural network architectures that operate of continuous and categorical input features which make hardware acceleration challenging. We will discuss the open research challenges and opportunities within this space. David Brooks 0001, Martin M. Frank, Tayfun Gokmen, Udit Gupta 0001, Xiaobo Sharon Hu, Shubham Jain 0004, Ann Franchesca Laguna, Michael T. Niemier, Ian O'Connor, Anand Raghunathan, Ashish Ranjan 0001, Dayane Reis, Jacob R. Stevens, Carole-Jean Wu, Xunzhao Yin |
DATE | 8 |
| 2020 | A Fast and Energy Efficient Computing-in-Memory Architecture for Few-Shot Learning ApplicationsabstractAmong few-shot learning methods, prototypical networks (PNs) are one of the most popular approaches due to their excellent classification accuracies and network simplicity. Test examples are classified based on their distances from class prototypes. Despite the application-level advantages of PNs, the latency of transferring data from memory to compute units is much higher than the PN computation time. Thus, PNs performance is limited by memory bandwidth. Computing-in-memory addresses this bandwidth-bottleneck problem by bringing a subset of compute units closer to memory. In this work, we propose a CiM-PN framework that enables the computation of distance metrics and prototypes inside the memory. CiM-PN replaces the computationally intensive Euclidean distance metric by the CiM-friendly Manhattan distance metric. Additionally, prototypes are computed using an in-memory mean operation realized by accumulation and division by powers of two, which enables few-shot learning implementations where "shots" are powers of two. The CiM-PN hardware uses CMOS memory cells, as well as CMOS peripherals such as customized sense amplifiers, carry-look-ahead adders, in-place copy buffers and a logarithmic shifter. Compared with a GPU implementation, a CMOS-based CiM-PN achieves speedups of 2808x/111x and energy savings of 2372x/5170x at iso-accuracy for the prototype and nearest-neighbor computation, respectively, and over 2x end-to-end speedup and energy improvements. We also gain 3-14% accuracy improvement when compared to existing non-GPU hardware approaches due to the floating-point CiM operations. Dayane Reis, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 3 |
| 2020 | A Novel TIGFET-based DFF Design for Improved Resilience to Power Side-Channel AttacksabstractSide-channel attacks (SCAs) represent a significant security threat, and aim to reveal otherwise secret data by analyzing a relevant circuit's behavior, e.g., its power consumption. While all circuit components are potential power side channels, D-flip-flops (DFFs) are often the primary source of information leakage to an SCA. This paper proposes a DFF design based on the three-independent-gate field-effect transistor (TTGFET) that reduces side-channel vulnerabilities of sequential circuits. Notably, we find that the I-V characteristics of the TIGFET itself leads to inherent side-channel resilience, which in turn enables simpler and more efficient cryptographic hardware. Our proposed design is based on a prior TIGFET-based true single-phase clock (TSPC) DFF design, which offers high performance and reduced area. More specifically, our modified TSPC (mTSPC) design exploits the symmetric I-V characteristics of TIGFETs, which results in pull-up and pull-down currents that are nearly identical. When combined with additional circuit modifications (made possible by the unique characteristics of the TIGFET), the mTSPC circuit draws almost the same amount of supply currents under all possible input transitions (less than 1% variation for different transitions), which can in turn mask information leakage. Using a 10nm TIGFET technology model, simulation results show that the proposed TIGFET-based DFF circuit leads to decreased power consumption (up to 96.9% when compared to the prior secured designs), has a low delay (15.2 ps), and employs only 12 TIGFET devices. Furthermore, an 8-bit S-box whose output is sampled by a group of eight mTSPC DFFs was simulated. A correlation power analysis attack on the simulated S-box with 256 power traces shows that the key is not revealed, which confirms the SCA resiliency of the proposed DFF design. Mohammad Mehdi Sharifi, Ramin Rajaei, Patsy Cadareanu, Pierre-Emmanuel Gaillardon, Yier Jin, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 6 |
| 2020 | AxR-NN: Approximate Computation Reuse for Energy-Efficient Convolutional Neural NetworksabstractThe recent success of convolutional neural networks (CNN) has led its implementation in specialized accelerators such as graphics processing unit (GPUs). However, the intensive computing workloads of CNNs remain a challenge to existing accelerators. By leveraging the error tolerance of CNNs, we propose a novel method to design energy-efficient CNN accelerators using approximate computation reuse (ACR), referred to as AxRNN. Computation reuse aims to reuse the previously computed results to avoid redundant executions. However, it cannot be applied directly to CNNs because CNNs do not have enough data locality. Thus, AxRNN performs approximate computation reuse under relaxed precision requirements on input patterns and design a reconfigurable architecture to support the ACR. This reconfigurable pattern matching is central to achieve a "controllable approximation". We implement the AxRNN using content addressable memory and integrate them with floating point units. Simulation results show that AxRNN reduces the computation energy by 30-58% with only 1-2.5% accuracy degradation on MNIST, EMNIST, and CIFAR-10 dataset. Dongning Ma, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu, Xun Jiao 0002 |
ACM Great Lakes Symposium on VLSI | 3 |
| 2020 | Modeling and Benchmarking Computing-in-Memory for Design Space ExplorationabstractThe bottleneck between the limited memory bandwidth and high speed processing demands is the main cause of problems associated with high volume of data transfers in data-intensive applications. As a possible remedy to these issues, computing-in-memory (CiM) enables a subset of logic and arithmetic operations to be performed where the data resides, i.e., inside the memory. Various CiM designs have been proposed to date, based on different technologies. Given the variety of options available, picking the right design option for a system/application can be a complex task. When choosing a CiM design, it is important to establish evaluation conditions that are as uniform as possible to make a fair choice between available design options. In this paper, we describe a methodology for an uniform benchmarking of CiM designs. Our approach evaluates devices/circuits, arrays and the overall impact of CiM to a system with a framework based on Eva-CiM. As a case study, we analyze the array-level performance of 7 recent CiM designs implemented with SRAM, DRAM, FeFET-RAM, STT-MRAM, SOT-MRAM, and RRAM. After we identify that the FeFET-RAM-based design shows promising energy and delay savings at the array level, we carry out a system level evaluation showing that FeFET-RAM-based CiM outperforms a CMOS SRAM CiM baseline by an average of 60% across a set of 17 benchmarks (with respect to energy savings). Regarding speedups, both technologies offer virtually the same benefit of about 1.5X when compared to a situation where processing does not happen in memory. Dayane Reis, Shaahin Angizi, Xunzhao Yin, Deliang Fan, Michael T. Niemier, Cheng Zhuo, Xiaobo Sharon Hu |
ACM Great Lakes Symposium on VLSI | 6 |
| 2020 | Seed-and-Vote based In-Memory Accelerator for DNA Read MappingabstractGenome analysis is becoming more important in the fields of forensic science, medicine, and history. Sequencing technologies such as High Throughput Sequencing (HTS) and Third Generation Sequencing (TGS) have greatly accelerated genome sequencing. However, genome read mapping remains significantly slower than sequencing. Because of the enormous amount of data needed, the speed of the data transfer between the memory and the processing unit limits the execution speed. In-memory computing can help address the memory-bandwidth bottleneck by minimizing data transfers. Ternary Content Addressable Memories (TCAMs) have been used in accelerators because of their fast searching capability for seed-and-extend, a popular read mapping approach. Seed-and-vote, another read mapping approach, is faster than the seed-and-extend approach but has lower accuracies when used with very short reads. Since sequencing technology is moving to longer reads, the seed-and-vote approach is becoming more viable. We propose a genome read mapping accelerator that uses approximate TCAM to execute the Fast Seed and Vote algorithm (FSVA) that can map both short and long reads. We achieved 400X acceleration compared to the seed-and-extend approach BWA-MEM on a CPU and 115X acceleration at 30X energy improvement compared to state-of-the-art in-memory accelerator using the seed-and-extend approach at 98.75% accuracy for 100bp reads. Ann Franchesca Laguna, Hasindu Gamaarachchi, Xunzhao Yin, Michael T. Niemier, Sri Parameswaran, Xiaobo Sharon Hu |
ICCAD | 4 |
| 2020 | A Hybrid FeMFET-CMOS Analog Synapse Circuit for Neural Network Training and InferenceabstractAn analog synapse circuit based on ferroelectric-metal field-effect transistors is proposed, that offers 6-bit weight precision. The circuit is comprised of volatile least significant bits (LSBs) used solely during training, and non-volatile most significant bits (MSBs) used for both training and inference. The design works at a 1.8V logic-compatible voltage, provides 1010endurance cycles, and requires only 250ps update pulses. A variant of LeNet trained with the proposed synapse achieves 98.2% accuracy on MNIST, which is only 0.4% lower than an ideal implementation of the same network with the same bit precision. Furthermore, the proposed synapse offers improvements of up to 26% in area, 44.8% in leakage power, 16.7% in LSB update pulse duration, and two orders of magnitude in endurance cycles, when compared to state-of-the-art hybrid synaptic circuits. Our proposed synapse can be extended to an 8-bit design, enabling a VGG-like network to achieve 88.8% accuracy on CIFAR-10 (only 0.8% lower than an ideal implementation of the same network). Arman Kazemi, Ramin Rajaei, Kai Ni 0004, Suman Datta, Michael T. Niemier, Xiaobo Sharon Hu |
ISCAS | 5 |
| 2020 | Dynamic Memory and Sequential Logic Design using Negative Capacitance FinFETsabstractThe emerging negative capacitance FinFET (NC-FinFET) device is a promising technology for the design of low-power VLSI circuits. This paper proposes ultra-low-power, highperformance, and low-area dynamic random access memory and sequential logic circuits based on NC-FinFETs. These circuits leverage the fact that NC-FinFETs have low leakage currents which help to facilitate a dynamic storage (DS)-based logic design style. This can in turn lead to reduced area overhead and help to compensate for higher delays that may be associated with NC-FinFETs. Our proposed circuit-level solutions can improve the data retention time of DS latch, flip-flop, and eDRAM circuits. Simulations with 14nm baseline FinFET (BS-FinFET) and NC-FinFET device models reveal that the proposed circuits offer up to 83.5% improvement in area-power-delay-product when compared to conventional BS-FinFET static-storage counterparts. Ramin Rajaei, Yen-Kai Lin, Sayeef S. Salahuddin, Michael T. Niemier, Xiaobo Sharon Hu |
ISCAS | 4 |
| 2020 | Embedding error correction into crossbars for reliable matrix vector multiplication using emerging devicesabstractEmerging memory devices are an attractive choice for implementing very energy-efficient in-situ matrix-vector multiplication (MVM) for use in intelligent edge platforms. Despite their great potential, device-level non-idealities have a large impact on the application-level accuracy of deep neural network (DNN) inference. We introduce a low-density parity-check code (LDPC) based approach to correct non-ideality induced errors encountered during in-situ MVM. We first encode the weights using error correcting codes (ECC), perform MVM on the encoded weights, and then decode the result after in-situ MVM. We show that partial encoding of weights can maintain DNN inference accuracy while minimizing the overhead of LDPC decoding. Within two iterations, our ECC method recovers 60% of the accuracy in MVM computations when 5% of underlying computations are error-prone. Compared to an alternative ECC method which uses arithmetic codes, using LDPC improves AlexNet classification accuracy by 0.8% at iso-energy. Similarly, at iso-energy, we demonstrate an improvement in CIFAR-10 classification accuracy of 54% with VGG-11 when compared to a strategy that uses 2× redundancy in weights. Further design space explorations demonstrate that we can leverage the resilience endowed by ECC to improve energy efficiency (by reducing operating voltage). A 3.3× energy efficiency improvement in DNN inference on CIFAR-10 dataset with VGG-11 is achieved at iso-accuracy. Qiuwen Lou, Tianqi Gao, Patrick Faley, Michael T. Niemier, Xiaobo Sharon Hu, Siddharth Joshi 0001 |
ISLPED | 4 |
| 2020 | GC-eDRAM design using hybrid FinFET/NC-FinFETabstractGain cell embedded DRAMs (GC-eDRAM) are a potential alternative for conventional static random access memories thanks to their attractive advantages such as high density, low-leakage, and two-ported operation. As CMOS technology nodes scale down, the design of GC-eDRAM at deeply scaled nanometer nodes becomes more challenging. Deeply-scale technology nodes suffer from high leakage currents and result in low data retention times (DRTs) for GC-eDRAMs. Negative capacitance FinFETs (NC-FinFETs) are a promising emerging device for ultra-low-power VLSI design. Due to the lower leakage currents, NC-FinFETs can facilitate GC-eDRAM design with higher DRTs. We show that though NC-FinFETs have lower OFF currents and higher ION/IOFF ratios, their ON current is lower than FinFETs by approximately 30%, which results in lower performance. To benefit from the potential power efficiencies and the high DRTs of NC-FinFETs without sacrificing performance, we propose hybrid FinFET/NC-FinFET configurations for some prior 2T, 3T, and 4T GC-eDRAM cells. Simulations based on a 14nm experimentally calibrated NC-FinFET model suggest that the hybrid designs offer up to 96.8% and 86.3% improvements in DRT and static power consumption, respectively, when compared to the FinFET implementation. They also offer up to 47% read delay improvement over the NC-FinFET design. We also study the voltage scaling effects on DRT and refresh-energy of the proposed GC-eDRAM cells. The associated simulation results reveal that, with different supply voltages, the proposed hybrid 4T GC-eDRAM cell offers up to 370× less refresh-energy when compared to the other designs. Ramin Rajaei, Yen-Kai Lin, Sayeef S. Salahuddin, Michael T. Niemier, Xiaobo Sharon Hu |
ISLPED | 4 |
| 2020 | Algorithmic Acceleration of B/FV-Like Somewhat Homomorphic Encryption for Compute-Enabled RAM
Jonathan Takeshita, Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu, Taeho Jung |
SAC | 4 |
| 2020 | SearcHD: A Memory-Centric Hyperdimensional Computing With Stochastic TrainingabstractBrain-inspired hyperdimensional (HD) computing emulates cognitive tasks by computing with long binary vectors-also know as hypervectors-as opposed to computing with numbers. However, we observed that in order to provide acceptable classification accuracy on practical applications, HD algorithms need to be trained and tested on nonbinary hypervectors. In this article, we propose SearcHD, a fully binarized HD computing algorithm with a fully binary training. SearcHD maps every data points to a high-dimensional space with binary elements. Instead of training an HD model with nonbinary elements, SearcHD implements a full binary training method which generates multiple binary hypervectors for each class. We also use the analog characteristic of nonvolatile memories (NVMs) to perform all encoding, training, and inference computations in memory. We evaluate the efficiency and accuracy of SearcHD on a wide range of classification applications. Our evaluation shows that SearcHD can provide on average 31.1× higher energy efficiency and 12.8× faster training as compared to the state-of-the-art HD computing algorithms. Mohsen Imani, Xunzhao Yin, John Messerly, Saransh Gupta, Michael T. Niemier, Xiaobo Sharon Hu, Tajana Rosing |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2020 | Computing-in-Memory for Performance and Energy-Efficient Homomorphic EncryptionabstractHomomorphic encryption (HE) allows direct computations on encrypted data. Despite numerous research efforts, the practicality of HE schemes remains to be demonstrated. In this regard, the enormous size of ciphertexts involved in HE computations degrades computational efficiency. Near-memory processing (NMP) and computing-in-memory (CiM)—paradigms where computation is done within the memory boundaries—represent architectural solutions for reducing latency and energy associated with data transfers in data-intensive applications, such as HE. This article introduces CiM-HE, a CiM architecture that can support operations for the Brakerski/Fan–Vercauteren (B/FV) scheme, a somewhat HE scheme for general computation. CiM-HE hardware consists of customized peripherals, such as sense amplifiers, adders, bit shifters, and sequencing circuits. The peripherals are based on CMOS technology and could support computations with memory cells of different technologies. Circuit-level simulations are used to evaluate our CiM-HE framework assuming a 6T-SRAM memory. We compare our CiM-HE implementation against: 1) two optimized CPU HE implementations and 2) a field-programmable gate array (FPGA)-based HE accelerator implementation. Compared with a CPU solution, CiM-HE obtains speedups between$4.6\times $and$9.1\times $and energy savings between$266.4\times $and$532.8\times $for homomorphic multiplications (the most expensive HE operation). Also, a set of four end-to-end tasks, i.e., mean, variance, linear regression, and inference, are up to$1.1\times $,$7.7\times $,$7.1\times $, and$7.5\times $faster (and$301.1\times $,$404.6\times $,$532.3\times $, and$532.8\times $more energy efficient). Compared with CPU-based HE in previous work, CiM-HE obtains$14.3\times $speedup and$> 2600\times $energy savings. Finally, our design offers$2.2\times $speedup with$88.1\times $energy savings compared with a state-of-the-art FPGA-based accelerator. Dayane Reis, Jonathan Takeshita, Taeho Jung, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2019 | Design of Hardware-Friendly Memory Enhanced Neural NetworksabstractNeural networks with external memories have been proven to minimize catastrophic forgetting, a major problem in applications such as lifelong and few-shot learning. However, such memory enhanced neural networks (MENNs) often require a large number of floating point-based cosine distance metric calculations to perform necessary attentional operations, which greatly increases energy consumption and hardware cost. This paper investigates other distance metrics in such neural networks in order to achieve more efficient hardware implementations in MENNs. We propose using content addressable memories (CAMs) to accelerate and simplify attentional operations. Our hardware friendly approach implements fixed point L∞distance calculations via ternary content addressable memories (TCAM) and fixed point L1and L2distance calculations on a general purpose graphical processing unit (GPGPU). As a representative example, a 32-bit floating point-based cosine distance MENN with M · D multiplications has a 99.06% accuracy for the Omniglot 5-way 5-shot classification task. Based on our approach, with just 4-bit fixed point precision, a L∞- L1distance hardware accuracy of 90.35% can be achieved with just 16 TCAM lookups and 16·D addition and subtraction operations. With 4-bit precision and a L∞-L2distance, hardware classification accuracies of 96.00% are possible. Hence, 16 TCAM lookups and 16·D multiplication operations are needed. Assuming the hardware memory has 512 entries, the number of multiplication operations is reduced by 32x versus the cosine distance approach. Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 2 |
| 2019 | An Energy Efficient Non-Volatile Flip-Flop based on CoMET TechnologyabstractAs we approach the limits of CMOS scaling, researchers are developing "beyond-CMOS" technologies to sustain the technological benefits associated with device scaling. Spin-tronic technologies have emerged as a promising beyond-CMOS technology due to their inherent benefits over CMOS such as high integration density, low leakage power, radiation hardness, and non-volatility. These benefits make spintronic devices an attractive successor to CMOS-especially for memory circuits. However, spintronic devices generally suffer from slower switching speeds and higher write energy, which limits their usability. In an effort to close the energy-delay gap between CMOS and spintronics, device concepts such as CoMET (Composite-Input Magnetoelectric-base Logic Technology) have been introduced, which collectively leverage material phenomena such as the spin-Hall effect and the magnetoelectric effect to enable fast, energy efficient device operation. In this work, we propose a non-volatile flip-flop (NVFF) based on CoMET technology that is capable of achieving up to two orders of magnitude less write energy than CMOS. This low write energy (≈2 aJ) makes our CoMET NVFF especially attractive to architectures that require frequent backup operations-e.g., for energy harvesting non-volatile processors. Robert Perricone, Zhaoxin Liang, Meghna G. Mankalale, Michael T. Niemier, Sachin S. Sapatnekar, Jianping Wang 0006, Xiaobo Sharon Hu |
DATE | 4 |
| 2019 | Ferroelectric FET Based In-Memory Computing for Few-Shot LearningabstractAs CMOS technology advances, the performance gap between the CPU and main memory has not improved. Furthermore, the hardware deployed for Internet of Things (IoT) applications need to process ever growing volumes of data, which can further exacerbate the "memory wall". Computing-in-memory (CiM) architectures, where logic and arithmetic operations are performed in memory, can significantly reduce energy and latency overheads associated with data transfer, and potentially alleviate processor-memory bottlenecks. In this paper, we consider the utility of ternary content addressable memory (TCAM) arrays and CiM arrays based on ferroelectric field effect transistors (FeFETs) to support emerging machine learning models that can learn new classes of data with significantly less training overhead - highly desirable in IoT applications. Architecturally, we use TCAM and CiM arrays to implement the external memory module in a memory enhanced neural network (MENN) - which can be used to minimize catastrophic forgetting - a major problem in applications such as lifelong and few-shot learning. As a representative example, we achieve 95.14% accuracy for a few-shot learning task with the Omniglot data set by using a combined L∞ infinity and L1 distance metric computed via a TCAM-CiM cascaded architecture (as opposed to 99.06% accuracy assuming a GPU backed by DRAM). While there is a slight drop in accuracy, the TCAM-CiM approach is 4.34X faster and 4.18X more energy efficient than a CMOS implementation for the same task. The ability of an FeFET to serve as both a compact logic and storage element helps to enable dense CiM and TCAM structures that drive the aforementioned improvements to application-level figures of merit (FOMs). Ann Franchesca Laguna, Xunzhao Yin, Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu |
ACM Great Lakes Symposium on VLSI | 4 |
| 2019 | The Impact of Emerging Technologies on Architectures and System-level Management: Invited PaperabstractThe goal of this work is to introduce and discuss different kinds of emerging technologies for logic circuitry and memory with respect to the key question of how they will impact future system-on-chip architectures and system-level management techniques. It is obvious that emerging technologies should have an impact there in order to fully exploit their technological advantages but also in order to deal with any disadvantages they might come with. In this special session paper, three promising emerging technologies are presented: (i) Negative Capacitance Field-Effect Transistor (NCFET) as a new CMOS technology with advantages primarily for low-power design, (ii) Ferroelectric FET (FeFET) as a non-volatile, area-efficient and low-power combined logic and memory as well as (iii) a Phase-Change Memory (PCM) and Resistive RAM (ReRAM) offering a large potential for tackling the memory wall problem in the von Neumann architecture. Our analysis demonstrates that not only new computing paradigms are promoted by these new technologies, it will also be seen that the trade-offs between the classical design parameters of low power, performance etc. will shift and hence emerging technologies will offer new Pareto points in the design space of future on-chip architectures. In that context, this work is unique as it bridges the gap between the technology side and system/architecture-level side to draw a vision of new technologies and their impact on architectures and system-level management. Jörg Henkel, Hussam Amrouch, Martin Rapp, Sami Salamin, Dayane Reis, Xunzhao Yin, Michael T. Niemier, Cheng Zhuo, Xiaobo Sharon Hu, Hsiang-Yun Cheng, Chia-Lin Yang |
ICCAD | 8 |
| 2019 | A Uniform Modeling Methodology for Benchmarking DNN AcceleratorsabstractDeep Neural Networks (DNNs) have achieved tremendous success in many application domains. Inspired by its success, specialized accelerators have been and continue to be developed to process DNN workloads in an energy-efficient manner. The design space for DNN accelerators can be extremely large since they can employ different datapaths, data mapping strategies, circuits, and device technologies. To explore the design space for developing DNN accelerators, it is important to quickly estimate the energy cost associated with an accelerator. This paper introduces a uniform modeling framework, Eva-DNN, to estimate the dynamic energy (a major component of total energy) consumed by a DNN accelerator. Specifically, we model the number of accesses and associated energy cost at different levels of memory and functional units. We derive a uniform expression that estimates the number of accesses as a function of the number of basic operations normalized by data reuse and activity factor of corresponding units. To model the energy cost of an individual functional unit operation, we employ a device-level benchmarking approach. Eva-DNN can accurately model energy contributions from device technology, circuits, architecture, data mapping strategy, and network. We applied our model on three accelerator architectures from the literature, namely: Eyeriss, ShiDianNao, and TrueNorth. Results suggest that Eva-DNN can accurately estimate energy contributions from different architectural units, achieving 4.5% to 8.0% of deviation from energy costs obtained from hardware measurements for different DNN workloads. Indranil Palit, Qiuwen Lou, Robert Perricone, Michael T. Niemier, Xiaobo Sharon Hu |
ICCAD | 4 |
| 2019 | A Mixed Signal Architecture for Convolutional Neural NetworksabstractDeep neural network (DNN) accelerators with improved energy and delay are desirable for meeting the requirements of hardware targeted for IoT and edge computing systems. Convolutional neural networks (CoNNs) belong to one of the most popular types of DNN architectures. This article presents the design and evaluation of an accelerator for CoNNs. The system-level architecture is based on mixed-signal, cellular neural networks (CeNNs). Specifically, we present (i) the implementation of different layers, including convolution, ReLU, and pooling, in a CoNN using CeNN, (ii) modified CoNN structures with CeNN-friendly layers to reduce computational overheads typically associated with a CoNN, (iii) a mixed-signal CeNN architecture that performs CoNN computations in the analog and mixed signal domain, and (iv) design space exploration that identifies what CeNN-based algorithm and architectural features fare best compared to existing algorithms and architectures when evaluated over common datasets—MNIST and CIFAR-10. Notably, the proposed approach can lead to 8.7× improvements in energy-delay product (EDP) per digit classification for the MNIST dataset at iso-accuracy when compared with the state-of-the-art DNN engine, while our approach could offer 4.3× improvements in EDP when compared to other network implementations for the CIFAR-10 dataset. Qiuwen Lou, Chenyun Pan, John McGuinness, András Horváth, Azad Naeemi, Michael T. Niemier, Xiaobo Sharon Hu |
ACM J. Emerg. Technol. Comput. Syst. | 6 |
| 2019 | Ferroelectric FETs-Based Nonvolatile Logic-in-Memory CircuitsabstractAmong the beyond-complementary metal-oxide- semiconductor (CMOS) devices being explored, ferroelectric field-effect transistors (FeFETs) are considered as one of the most promising. FeFETs are being studied by all major semiconductor manufacturers, and experimentally, FeFETs are making rapid progress. FeFETs also stand out with the unique hysteretic Ids-Vgs characteristic that allows a device to function as both a switch and a nonvolatile (NV) storage element. We exploit this FeFET property to build two categories of fine-grained logic-in-memory (LiM) circuits: 1) ternary content addressable memory (TCAM) which integrates efficient and compact logic/processing elements into various levels of memory hierarchy; 2) basic logic function units for constructing larger and more complex LiM circuits. Two writing schemes (with and without negative supply voltages respectively) for FeFETs are introduced in our LiM designs. The resulting designs are compared with existing LiM approaches based on CMOS, magnetic tunnel junctions (MTJs), resistive random access memories (ReRAMs), ferrorelectric tunnel junctions (FTJs), etc., that afford the same circuit-level functionality. Simulation results show that FeFET-based NV TCAMs offer lower area overhead than MTJ (79%) and CMOS (42% less) equivalents, as well as better search energy-delay products (EDPs) than TCAM designs based on MTJ (149×), ReRAM (1.7×), and CMOS (1.3×) in array evaluations. NV FeFET-based LiM basic circuit blocks are also more efficient than functional equivalents based on MTJs in terms of propagation delay (4.2×) and dynamic power (2.5×). A case study for an FeFET-based LiM accumulator further demonstrates that by employing FeFET as both a switch and an NV storage element, the FeFET-based accumulator can save area (36%) and power consumption (40%) when compared with a conventional CMOS accumulator with the same structure. Xunzhao Yin, Xiaoming Chen 0003, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2018 | Biomedical Image Segmentation Using Fully Convolutional Networks on TrueNorthabstractWith the rapid growth of medical and biomedical image data, energy-efficient solutions for analyzing such image data that can be processed fast and accurately on platforms with low power budget are highly desirable. This paper uses segmenting glial cells in brain microscopy images as a case study to demonstrate how to achieve biomedical image segmentation with significant energy saving and minimal comprise in accuracy. Specifically, we design, train, implement, and evaluate Fully Convolutional Networks (FCNs) for biomedical image segmentation on IBM's neurosynaptic DNN processor - TrueNorth (TN). Comparisons in terms of accuracy and energy dissipation of TN with that of a low power NVIDIA TX2 mobile GPU platform have been conducted. Experimental results show that TN can offer at least two orders of magnitude improvement in energy efficiency when compared to TX2 GPU for the same workload. Indranil Palit, Lin Yang 0003, Yue Ma 0001, Danny Ziyi Chen, Michael T. Niemier, Jinjun Xiong, Xiaobo Sharon Hu |
CBMS | 5 |
| 2018 | Computing with ferroelectric FETs: Devices, models, systems, and applicationsabstractIn this paper, we consider devices, circuits, and systems comprised of transistors with integrated ferroelectrics. Said structures are actively being considered by various semiconductor manufacturers as they can address a large and unique design space. Transistors with integrated ferroelectrics could (i) enable a better switch (i.e., offer steeper subthreshold swings), (ii) are CMOS compatible, (iii) have multiple operating modes (i.e., I-V characteristics can also enable compact, 1-transistor, non-volatile storage elements, as well as analog synaptic behavior), and (iv) have been experimentally demonstrated (i.e., with respect to all of the aforementioned operating modes). These device-level characteristics offer unique opportunities at the circuit, architectural, and system-level, and are considered here from device, circuit/architecture, and foundry-level perspectives. Ahmedullah Aziz, Evelyn T. Breyer, Xiaoming Chen 0003, Suman Datta, Sumeet Kumar Gupta, Michael Hoffmann 0008, Xiaobo Sharon Hu, Adrian M. Ionescu, Matthew Jerry, Thomas Mikolajick, Halid Mulaosmanovic, Kai Ni 0004, Michael T. Niemier, Ian O'Connor, Atanu Saha, Stefan Slesazeck, Sandeep Krishna Thirumala, Xunzhao Yin |
DATE | 14 |
| 2018 | Design and optimization of FeFET-based crossbars for binary convolution neural networksabstractBinary convolution neural networks (CNNs) have attracted much attention for embedded applications due to low hardware cost and acceptable accuracy. Nonvolatile, resistive random-access memories (RRAMs) have been adopted to build crossbar accelerators for binary CNNs. However, RRAMs still face fundamental challenges such as sneak paths, high write energy, etc. We exploit another emerging nonvolatile device-ferroelectric field-effect transistor (FeFET), to build crossbars to improve the energy efficiency for binary CNNs. Due to the three-terminal transistor structure, an FeFET can function as both a nonvolatile storage element and a controllable switch, such that both write and read power can be reduced. Simulation results demonstrate that compared with two RRAM-based crossbar structures, our FeFET-based design improves write power by 5600× and 3950×, and read power by 4.1× and 3.1×. We also tackle an important challenge in crossbar-based CNN accelerators: when a crossbar array is not large enough to hold the weights of one convolution layer, how do we partition the workload and map computations to the crossbar array? We introduce a hardware-software co-optimization solution for this problem that is universal for any crossbar accelerators. Xiaoming Chen 0003, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 3 |
| 2018 | Nonvolatile Lookup Table Design Based on Ferroelectric Field-Effect TransistorsabstractAs a nonvolatile (NV) device, ferroelectric field-effect transistors (FeFETs) have the potential to reduced power and area by integrating NV storage elements into logic. In this paper, we exploit FeFET nonvolatility to design lookup tables (LUTs), which have obvious utility in field-programmable gate arrays, etc. With nonvolatility, a single FeFET can be used as a storage cell in an LUT, which can help reduce both power and area. We design both static and dynamic logic style LUTs. Read and write schemes are also designed for the proposed LUTs. Evaluation results show that our LUTs outperform both conventional static random-access memory based LUTs as well as other NV LUTs in term of area-power-delay product. Xiaoming Chen 0003, Michael T. Niemier, Xiaobo Sharon Hu |
ISCAS | 2 |
| 2018 | Computing in memory with FeFETsabstractData transfer between a processor and memory frequently represents a bottleneck with respect to improving application-level performance. Computing in memory (CiM), where logic and arithmetic operations are performed in memory, could significantly reduce both energy consumption and computational overheads associated with data transfer. Compact, low-power, and fast CiM designs could ultimately lead to improved application-level performance. This paper introduces a CiM architecture based on ferroelectric field effect transistors (FeFETs). The CiM design can serve as a general purpose, random access memory (RAM), and can also perform Boolean operations ((N)AND, (N)OR, X(N)OR, INV) as well as addition (ADD) between words in memory. Unlike existing CiM designs based on other emerging technologies, FeFET-CiM accomplishes the aforementioned operations via a single current reference in the sense amplifier, which leads to more compact designs and lower power. Furthermore, the high Ion/Ioff ratio of FeFETs enables an inexpensive voltage-based sense scheme. Simulation-based case studies suggest that our FeFET-CiM can achieve speed-ups (and energy reduction) of ~119X (~1.6X) and ~1.97X (~1.5X) over ReRAM and STT-RAM CiM designs with respect to in-memory addition of 32-bit words. Furthermore, our approach offers an average speedup of ~2.5X and energy reduction of ~1.7X when compared to a conventional (not in-memory) approach across a wide range of benchmarks. Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu |
ISLPED | 2 |
| 2018 | Cross-layer efforts for energy-efficient computing: towards peta operations per second per wattabstractAs Moore’s law based device scaling and accompanying performance scaling trends are slowing down, there is increasing interest in new technologies and computational models for fast and more energy-efficient information processing. Meanwhile, there is growing evidence that, with respect to traditional Boolean circuits and von Neumann processors, it will be challenging for beyond-CMOS devices to compete with the CMOS technology. Exploiting unique characteristics of emerging devices, especially in the context of alternative circuit and architectural paradigms, has the potential to offer orders of magnitude improvement in terms of power, performance, and capability. To take full advantage of beyond-CMOS devices, cross-layer efforts spanning from devices to circuits to architectures to algorithms are indispensable. This study examines energy-efficient neural network accelerators for embedded applications in this context. Several deep neural network accelerator designs based on cross-layer efforts spanning from alternative device technologies, circuit styles, to architectures are highlighted. Application-level benchmarking studies are presented. The discussions demonstrate that cross-layer efforts indeed can lead to orders of magnitude gain towards achieving extreme-scale energy-efficient processing. Xiaobo Sharon Hu, Michael T. Niemier |
Frontiers Inf. Technol. Electron. Eng. | 2 |
| 2017 | In Quest of the Next Information Processing Substrate: Extended Abstract: InvitedabstractConventional CMOS scaling and the Moore's law have been the cornerstone of progress in computing hardware technology. However, with dimensional scaling expected to end soon, there is a pressing need to find the next information processing hardware that can continue to support the technology revolution. Will this hardware solution be an enhanced or an augmented version of MOSFET or a switch based on a radically new switching mechanism. Ultimately, do we require a complete deviation from the Boolean paradigm itself? In this invited paper, we will review some of the actively pursued future logic, merged logic-memory and related concepts. Suman Datta, Alan C. Seabaugh, Michael T. Niemier, Arijit Raychowdhury, Darrell Schlom, Debdeep Jena, Huili Grace Xing, H.-S. Philip Wong, Eric Pop, Sayeef S. Salahuddin, Sumeet Kumar Gupta, Supratik Guha |
DAC | 3 |
| 2017 | A Pathway to Enable Exponential Scaling for the Beyond-CMOS Era: InvitedabstractMany key technologies of our society, including so-called artificial intelligence (AI) and big data, have been enabled by the invention of transistor and its ever-decreasing size and ever-increasing integration at a large scale. However, conventional technologies are confronted with a clear scaling limit. Many recently proposed advanced transistor concepts are also facing an uphill battle in the lab because of necessary performance tradeoffs and limited scaling potential. We argue for a new pathway that could enable exponential scaling for multiple generations. This pathway involves layering multiple technologies that enable new functions beyond those available from conventional and newly proposed transistors. The key principles for this new pathway have been demonstrated through an interdisciplinary team effort at C-SPIN (a STARnet center), where systems designers, device builders, materials scientists and physicists have all worked under one umbrella to overcome key technology barriers. This paper reviews several successful outcomes from this effort on topics such as the spin memory, logic-in-memory, cognitive computing, stochastic and probabilistic computing and reconfigurable information processing. Jianping Wang 0006, Sachin S. Sapatnekar, Chris H. Kim, Paul A. Crowell, Steven J. Koester, Supriyo Datta, Kaushik Roy 0001, Anand Raghunathan, Xiaobo Sharon Hu, Michael T. Niemier, Azad Naeemi, Chia-Ling Chien, Caroline A. Ross, Roland Kawakami |
DAC | 10 |
| 2017 | Cellular neural network friendly convolutional neural networks - CNNs with CNNsabstractThis paper discusses the development and evaluation of a Cellular Neural Network (CeNN) friendly deep learning network for solving the MNIST digit recognition problem. Prior work has shown that CeNNs leveraging emerging technologies such as tunnel transistors can improve energy or EDP of CeNNs, while simultaneously offering richer/more complex functionality. Important questions to address are what applications can benefit from CeNNs, and whether CeNNs can eventually outperform other alternatives at the application-level in terms of energy, performance, and accuracy. This paper begins to address these questions by using the MNIST problem as a case study. András Horváth, Michael Hillmer, Qiuwen Lou, Xiaobo Sharon Hu, Michael T. Niemier |
DATE | 5 |
| 2017 | Advanced spintronic memory and logic for non-volatile processorsabstractMany ultra-low power Internet of things (IoT) systems may be powered by energy harvested from ambient sources (e.g., solar radiation, thermal gradients, and WiFi). However, these energy sources can vary significantly in terms of their strengths and on/off patterns. For volatile systems, the intermittent nature of the energy sources necessitates the use of backup/recovery schemes to guarantee computational correctness and forward progress, which incur performance, area and energy overhead. Non-volatile (NV) processors based on spintronic devices, such as Spin-Transfer Torque (STT) memory and All-Spin-Logic (ASL), are more attractive alternatives. These NV devices are capable of achieving forward progress without relying on backup/recovery schemes. This work establishes a general framework for evaluating NV device-based processors for energy harvesting applications. Results demonstrate that NV spintronic processors can achieve significant energy savings (up to 83 x) versus a hybrid CMOS (computation) and STT-RAM (backup) implementation. Robert Perricone, Ibrahim Ahmed 0002, Zhaoxin Liang, Meghna G. Mankalale, Xiaobo Sharon Hu, Chris H. Kim, Michael T. Niemier, Sachin S. Sapatnekar, Jianping Wang 0006 |
DATE | 7 |
| 2017 | Design and benchmarking of ferroelectric FET based TCAMabstractWe consider how emerging transistor technologies, specifically ferroelectric field effect transistors (or FeFETs), can realize compact and energy efficient ternary content addressable memories (TCAMs). As Moore's Law-based performance scaling trends slow, and many computational tasks of interest are now more data-centric than compute-centric, researchers are looking to improve performance/save energy by integrating efficient and compact logic/processing elements into various levels of the memory hierarchy. Potential benefits include reduced I/O traffic, energy/delay from data transfers, etc. A TCAM is an example of a logic-in-memory element that is ubiquitous in routers, caches, databases, and even neural networks. Not surprisingly, researchers continue to study how emerging technologies could lead to improved TCAMs. Recent work has considered how non-volatile (NV) memory technologies (e.g., resistive random access memory (ReRAM) or magnetic tunnel junctions (MTJs)) could best be used to construct low energy, NV TCAMs. However, acceptable Ron-Roffratios and the two terminal nature of these devices introduce energy and area overheads. Due to hysteresis in a device's I-V curve, an FeFET-based NV TCAM, offers low area overhead, as well as search energies and search speeds that are superior to other TCAM designs (i.e., based on MTJ, ReRAM and CMOS in array- and architectural-level evaluations). Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 2 |
| 2017 | Exploiting Non-Volatility for Information ProcessingabstractThe emergence of non-volatile (NV) technologies provides an opportunity to overcome limitations of CMOS (e.g., the growth of leakage power), while simultaneously providing a degree of non-volatility to the processor. Spintronic NV technologies are of particular interest due to their high integration density, low device count, radiation hardness, and non-volatility when compared to CMOS. Quantifying the impact of spintronic NV technologies at the architecture/application levels introduces a unique challenge as the granularity of technology integration can vary significantly (i.e., from heterogeneous architectures with NV cache/memory and CMOS-based logic, to completely NV architectures). In this work, we explore three classes of NV processors (NVPs) and define metrics for quantifying their respective energy savings. As case studies, we evaluate the impact of NV technologies for both an energy harvesting non-pipelined processor and a general purpose processor executing scientific applications under varying degrees of parallelism. Robert Perricone, Li Tang 0007, Michael T. Niemier, Xiaobo Sharon Hu |
ACM Great Lakes Symposium on VLSI | 3 |
| 2016 | Using emerging technologies for hardware security beyond PUFs
Xiaobo Sharon Hu, Yier Jin, Michael T. Niemier, Xunzhao Yin |
DATE | 4 |
| 2016 | Can beyond-CMOS devices illuminate dark silicon?
Robert Perricone, Xiaobo Sharon Hu, Joseph Nahas, Michael T. Niemier |
DATE | 4 |
| 2016 | Design of latches and flip-flops using emerging tunneling devices
Xunzhao Yin, Behnam Sedighi, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 3 |
| 2016 | Enhancing Hardware Security with Emerging Transistor TechnologiesabstractWe consider how the I-V characteristics of emerging transistors (particularly those sponsored by STARnet) might be employed to enhance hardware security. An emphasis of this work is to move beyond hardware implementations of physically unclonable functions (PUFs) and random num- ber generators (RNGs). We highlight how new devices (i) may enable more sophisticated logic obfuscation for IP protection, (ii) could help to prevent fault injection attacks, (iii) prevent differential power analysis in lightweight cryptographic systems, etc. Yu Bi, Xiaobo Sharon Hu, Yier Jin, Michael T. Niemier, Kaveh Shamsi, Xunzhao Yin |
ACM Great Lakes Symposium on VLSI | 4 |
| 2016 | Exploiting ferroelectric FETs for low-power non-volatile logic-in-memory circuitsabstractNumerous research efforts are targeting new devices that could continue performance scaling trends associated with Moore's Law and/or accomplish computational tasks with less energy. One such device is the ferroelectric FET (FeFET), which offers the potential to be scaled beyond the end of the silicon roadmap as predicted by ITRS. Furthermore, the Ids vs. Vgs characteristics of FeFETs may allow a device to function as both a switch and a non-volatile storage element. We exploit this FeFET property to enable fine-grained logic-in-memory (LiM). We consider three different circuit design styles for FeFET-based LiM: complementary (differential), dynamic current mode, and dynamic logic. Our designs are compared with existing approaches for LiM (i.e., based on magnetic tunnel junctions (MTJs), CMOS, etc.) that afford the same circuit-level functionality. Assuming similar feature sizes, non-volatile FeFET-based LiM circuits are more efficient than functional equivalents based on MTJs when considering metrics such as propagation delay (2.9×, 6.8×) and dyanmic power (3.7×, 2.3×) (for 45 nm, 22 nm technology respectively). Compared to CMOS functional equivalents, FeFET designs still exhibit modest improvements in the aforementioned metrics while also offering non-volatility and reduced device count. Xunzhao Yin, Ahmedullah Aziz, Joseph Nahas, Suman Datta, Sumeet Kumar Gupta, Michael T. Niemier, Xiaobo Sharon Hu |
ICCAD | 6 |
| 2016 | Emerging Technology-Based Design of Primitives for Hardware SecurityabstractHardware security concerns such as intellectual property (IP) piracy and hardware Trojans have triggered research into circuit protection and malicious logic detection from various design perspectives. In this article, emerging technologies are investigated by leveraging their unique properties for applications in the hardware security domain. Security, for the first time, will be treated as one design metric for emerging nano-architecture. Five example circuit structures including camouflaging gates, polymorphic gates, current/voltage-based circuit protectors, and current-based XOR logic are designed to show the high efficiency of silicon nanowire FETs and graphene SymFET in applications such as circuit protection and IP piracy prevention. Simulation results indicate that highly efficient and secure circuit structures can be achieved via the use of non-CMOS devices. Yu Bi, Kaveh Shamsi, Jiann-Shiun Yuan, Pierre-Emmanuel Gaillardon, Giovanni De Micheli, Xunzhao Yin, Xiaobo Sharon Hu, Michael T. Niemier, Yier Jin |
ACM J. Emerg. Technol. Comput. Syst. | 8 |
| 2015 | Towards systematic design of 3D pNML layouts
Robert Perricone, Yining Zhu, Katherine M. Sanders, Xiaobo Sharon Hu, Michael T. Niemier |
DATE | 5 |
| 2015 | A CNN-inspired mixed signal processor based on tunnel transistors
Behnam Sedighi, Indranil Palit, Xiaobo Sharon Hu, Joseph Nahas, Michael T. Niemier |
DATE | 5 |
| 2015 | TFET-based Operational Transconductance Amplifier Design for CNN SystemsabstractA Cellular Neural Network (CNN) is a powerful processor that can significantly improve the performance of spatio-temporal applications such as pattern recognition, image processing, motion detection, when compared to the more traditional von Neumann architecture. In this paper, we show how tunneling field effect transistors (TFETs) can be utilized to enhance the performance of CNNs. Specifically, power consumption of TFET-based CNNs can be significantly lower when compared to MOSFET-based CNNs due to improved voltage controlled current sources (VCCSs) - an important component in CNN systems. We demonstrate that CNNs can benefit from low power conventional linear VCCSs implemented via TFETs. We also show that TFETs can be useful to realize non-linear VCCSs, which are either not possible or exhibit degraded performance when implemented via CMOS. Such non-linear VCCSs help to improve the performance of certain CNN operations (e.g., global maximum/minimum). We provide two case studies - image contrast enhancement and maximum row selection - that illustrate the benefits of non-linear VCCSs (e.g., reduced computation time, energy dissipation, etc.) when compared to CMOS-based approaches. Qiuwen Lou, Indranil Palit, András Horváth, Xiaobo Sharon Hu, Michael T. Niemier, Joseph Nahas |
ACM Great Lakes Symposium on VLSI | 5 |
| 2015 | Analytically Modeling Power and Performance of a CNN SystemabstractCellular neural networks (CNNs) are a powerful analog architecture that can outperform traditional von Neumann architecture for spatio-temporal information processing applications, e.g., image processing and speech recognition. Much existing work reports energy dissipation for CNNs at the chip level, which includes dissipation of sensors, actuators, and other components. As such, the impacts of various system variables, e.g., application templates, characteristics of the resistive element, etc., on the energy profile of a CNN cannot be easily determined. In this work, we propose analytical models to estimate CNN power and performance (measured by settling time). Power dissipations, and settling times obtained via the models for different linear, and non-linear characteristics are verified through circuit simulation. Simulation results show that the proposed models predict power dissipation and settling time with less than 1% and 3% errors, respectively. By using these models, we have also performed case studies for a tactile sensing problem, and a pattern recognition problem to compare power and performance between tunneling field effect transistor (TFET) based non-linear CNN and conventional linear resistor based CNN. Indranil Palit, Qiuwen Lou, Nicholas Acampora, Joseph Nahas, Michael T. Niemier, Xiaobo Sharon Hu |
ICCAD | 5 |
| 2015 | Reliable and high performance STT-MRAM architectures based on controllable-polarity devicesabstractSource degeneration of access devices in the parallel (P)_ anti-parallel (AP) switching in Spin Transfer Torque Magnetic Random Access Memories (STT-MRAM) has ultimately been a limiting factor in the operational speed of these types of memories. In this work, new architectures for memory single-cells and arrays of cells are presented that utilize Schottky-Barrier Silicon Nanowire Field Effect Transistors with polarity control capabilities (e.g., SiNW-FETs), to substantially increase the performance of STT-MRAM, specifically Multi-Level Cell (MLC) STT-MRAM. The proposed design offers built-in reliability improvement as it omits one of the available four states in the MLC STT-MRAM memory facilitating the resistance level detection for peripheral circuitry. Our simulation results of the developed memory cell show 49.7% reductions in P-AP switching time, as well as 51.3% increases in available drive current under 1.4V supply voltage when compared to FinFET 22imi technology. With respect to memory arrays, the proposed architecture demonstrates an average write latency reduction of 37% in comparison with FinFET 22nm technology node. Kaveh Shamsi, Yu Bi, Yier Jin, Pierre-Emmanuel Gaillardon, Michael T. Niemier, Xiaobo Sharon Hu |
ICCD | 5 |
| 2014 | Leveraging Emerging Technology for Hardware Security - Case Study on Silicon Nanowire FETs and Graphene SymFETsabstractHardware security concerns such as IP piracy and hardware Trojans have triggered research into circuit protection and malicious logic detection from various design perspectives. In this paper, emerging technologies are investigated by leveraging their unique properties for applications in the hardware security domain. Three example circuit structures including camouflaging gates, polymorphic gates and power regulators are designed to prove the high efficiency of silicon nanowire FETs and graphene Sym FET in applications such as circuit protection and IP piracy prevention. Simulation results indicate that highly efficient and secure circuit structures can be achieved via the use of emerging technologies. Yu Bi, Pierre-Emmanuel Gaillardon, Xiaobo Sharon Hu, Michael T. Niemier, Jiann-Shiun Yuan, Yier Jin |
ATS | 4 |
| 2014 | Impact of steep-slope transistors on non-von Neumann architectures: CNN case studyabstractA Cellular Neural Network (CNN) is a highly-parallel, analog processor that can significantly outperform von Neumann architectures for certain classes of problems. Here, we show how emerging, beyond-CMOS devices could help to further enhance the capabilities of CNNs, particularly for solving problems with non-binary outputs. We show how CNNs based on devices such as graphene transistors - with multiple steep current growth regions separated by negative differential resistance (NDR) in their I-V characteristics - could be used to recognize multiple patterns simultaneously. (This would require multiple steps given a conventional, binary CNN.) Also, we demonstrate how tunneling field effect transistors (TFETs) can be used to form circuits capable of performing similar tasks. With this approach, more “exotic” device I-V characteristics are not required - which should be an asset when considering issues such as cell-to-cell mismatch, etc. As a case study, we present a CNN-cell design that employs TFET-based circuitry to realize ternary outputs. We then illustrate how this hardware could be employed to efficiently solve a tactile sensing problem. The total number of computation steps as well as the required hardware could be reduced significantly when compared to an approach based on a conventional CNN. Indranil Palit, Behnam Sedighi, András Horváth, Xiaobo Sharon Hu, Joseph Nahas, Michael T. Niemier |
DATE | 6 |
| 2014 | Design of 3D nanomagnetic logic circuits: A full-adder case studyabstractNanomagnetic logic (NML) is a “beyond-CMOS” technology that combines logic and memory capabilities through field-coupled interactions between nanoscale magnets. NML is intrinsically non-volatile, low-power, and radiation-hard when compared to CMOS equivalents. Moreover, there have been numerous demonstrations of NML circuit functionality within the last decade. These fabricated structures typically employ devices with in-plane magnetization to move and process data. However, in-plane layouts imply circuits and interconnects in only two dimensions (2D), which makes signal routing - and hence circuits - more complex. In this paper, we introduce NML circuits that move and process data in three dimensions (3D). We employ devices with perpendicular magnetic anisotropy (PMA) (i.e., out-of-plane magnetization states) and discuss their behavior when utilized in 3D designs. Furthermore, we provide a systematic design approach for 3D NML circuits using a threshold full adder as a case study. We compare our 3D adder to 2D adders to highlight the benefits of 3D NML circuits, which include simpler signal routing and a smaller area footprint. Robert Perricone, Xiaobo Sharon Hu, Joseph Nahas, Michael T. Niemier |
DATE | 4 |
| 2014 | Cellular neural networks for image analysis using steep slope devicesabstractTraditional CMOS based von Neumann architectures face daunting challenges in performing complex computational tasks at high speed and with low power on spatio-temporal data, e.g., image processing, pattern recognition, etc. In this study, we discuss the utilities of various steep slope, beyond-CMOS emerging devices for image processing applications within the non-von Neumann computing paradigm of cellular neural networks (CNNs). In general, the steep subthreshold swing of the devices obviates the output transfer hardware used in a conventional CNN cell. For image processing with binary stable outputs, Tunnelling FETs (TFETs) can facilitate low power operation. For multi-valued problems, devices like graphene transistors, Symmetric tunnelling FETs (SymFETs) might be leveraged to solve a problem with fewer computational steps. The potential for additional hardware reduction when compared to functional equivalents via conventional CNNs is also possible. Emerging devices can also lead to lower power implementations of the voltage controlled current sources (VCCSs) that are an integral component of any CNN cell. Furthermore, non-linear implementations of the VCCSs via emerging devices could enable simpler computational paths for many image processing tasks. Indranil Palit, Qiuwen Lou, Michael T. Niemier, Behnam Sedighi, Joseph Nahas, Xiaobo Sharon Hu |
ICCAD | 3 |
| 2014 | Boolean circuit design using emerging tunneling devicesabstractNovel device technologies are exceedingly under investigation for the sub-10-nm era. Some tunneling devices employing 2-D materials have shown the potential for low-voltage operation, promising energy efficient digital circuits and systems. Interestingly, certain emerging tunneling devices such as SymFETs and BiSFETs exhibit an I-V characteristic different from that of MOSFETs. In this paper, the design of Boolean gates with SymFETs is studied. We show that the negative differential resistance (NDR) behavior of the transistors leads to hysteresis in inverters and buffers, and can be used to build simple Schmitttriggers. It can also by used in designing new pseudo-SymFET loads for circuits similar to all-n-type or dynamic logic. We demonstrate the feasibility of building NAND, NOR, IMPLY, and MAJORITY gates with fewer transistors when compared with static CMOS designs. Benchmarking efforts show that SymFETs are an attractive choice for applications that demand low power and have moderate speed requirements, and demonstrate better dynamic energy efficiency than CMOS circuits; but one challenge for SymFET circuits is relatively larger leakage currents. Behnam Sedighi, Joseph Nahas, Michael T. Niemier, Xiaobo Sharon Hu |
ICCD | 3 |
| 2013 | GPU acceleration of Data Assembly in Finite Element Methods and its energy implicationsabstractThe Finite Element Method (FEM) is a numerical technique widely used in finding approximate solutions for many scientific and engineering problems. The Data Assembly (DA) stage in FEM can take up to 50% of the total FEM execution time. Accelerating DA with Graphics Processing Units (GPUs) presents challenges due to DA's mixed compute-intensive and memory-intensive workloads. This paper uses a representative finite element mini-application to explore DA acceleration on CPU+GPU platforms. Implementations based on different thread, kernel and task design approaches are developed and compared. Their performance and energy consumption are measured on four CPU+GPU and two CPU only platforms. The results show that (i) the performance and energy for different implementations on the same platform can vary significantly but the performance and energy trends are the same, and (ii) there exist performance and energy tradeoffs across some platforms if the best implementation is chosen for each of the platforms. Li Tang 0007, Xiaobo Sharon Hu, Danny Ziyi Chen, Michael T. Niemier, Richard F. Barrett, Simon D. Hammond, Genie Hsieh |
ASAP | 4 |
| 2013 | Minimum-energy state guided physical design for nanomagnet logicabstractNanomagnet Logic (NML) accomplishes computation through magnetic dipole-dipole interactions. It has the potential for low-power dissipation, radiation hardness and non-volatility. NML circuits have been designed to process and move information via nearest neighbor, device-to-device coupling. However, the resultant layouts often fail to function correctly. This paper reveals an important cause of such failures showing that a robust NML layout must take into account not only nearest neighbor, but also the next nearest neighbor couplings. A new design method is then introduced to address this issue that leverages the minimum-energy states of an NML circuit to guide the layout process. Case studies show that the new method is efficient and effective in arriving at correct NML layouts. Shiliang Liu, György Csaba, Xiaobo Sharon Hu, Edit Varga, Michael T. Niemier, Gary H. Bernstein, Wolfgang Porod |
DAC | 5 |
| 2013 | Systematic design of nanomagnet logic circuitsabstractNanomagnet Logic (NML) is an emerging device architecture that performs logic operations through fringing field interactions between nano-scale magnets. The design space for NML circuits is large and so far there exists no systematic approach for determining the parameter values (e.g., device-to-device spacings, clocking field strength etc.) to generate a predictable design solution. This paper presents a formal methodology for designing NML circuits that marshals the design parameters to generate a layout that is guaranteed to evolve correctly in time at 0K. The approach is further augmented to identify functional design targets when considering thermal noise associated with higher temperatures. The approach is applied to identify layouts for a 2-input AND gate, a “corner turn,” and a 3-input majority gate. Layouts are verified through simulations both at 0K and room temperature (300K). Indranil Palit, Xiaobo Sharon Hu, Joseph Nahas, Michael T. Niemier |
DATE | 4 |
| 2013 | TFET-based cellular neural network architecturesabstractIt is well known that CMOS scaling trends are now accompanied by less desirable byproducts such as increased energy dissipation. To combat the aforementioned challenges, solutions are sought at both the device and architectural levels. With this context, this work focuses on embedding a low voltage device, a Tunneling Field Effect Transistor (TFET) within a Cellular Neural Network (CNN) - a low power analog computing architecture. Our study shows that TFET-based CNN systems, aside from being fully functional, also provide significant power savings when compared to the conventional resistor-based CNN. Our initial studies suggest that power savings are possible by carefully engineering lower voltage, lower current TFET devices without sacrificing performance. Moreover, TFET-based CNN reduces implementation footprints by eliminating the hardware required to realize output transfer functions. Application dynamics are verified through simulations. We conclude the paper with a discussion of desired device characteristics for CNN architectures with enhanced functionality. Indranil Palit, Xiaobo Sharon Hu, Joseph Nahas, Michael T. Niemier |
ISLPED | 4 |
| 2012 | Making non-volatile nanomagnet logic non-volatileabstractField-coupled nanomagnets can offer significant energy savings at iso-performance versus CMOS equivalents. Magnetic logic could be integrated with CMOS, operate in environments that CMOS cannot, and retain state without power. Clocking requirements lead to inherently pipelined circuits, and high throughput further improves application-level performance. However, bit conflicts -- that will occur in defect free, pipelined ensembles -- can make non-volatile logic volatile. Assuming a field-based clock, we present hardware designs to improve steady state non-volatility, and explain how design enhancements could increase clock energy. We then suggest materials-related design levers that could simultaneously deliver non-volatility and low clock energy. Aaron Dingler, Steve Kurtz, Michael T. Niemier, Xiaobo Sharon Hu, György Csaba, Joseph Nahas, Wolfgang Porod, Gary H. Bernstein, Vjiay Karthik Sankar |
DAC | 3 |
| 2012 | A Reconfigurable PLA Architecture for Nanomagnet LogicabstractIn order to continue the performance and scaling trends that we have come to expect from Moore’s Law, many emergent computational models, devices, and technologies are actively being studied to either replace or augment CMOS technology. Nanomagnet Logic (NML) is one such alternative. NML operates at room temperature, it has the potential for low power consumption, and it is CMOS compatible. In this aricle, we present an NML programmable logic array (PLA) based on a previously proposed reprogrammable quantum-dot cellular automata PLA design. We also discuss the fabrication and simulation validation of the circuit structures unique to the NML PLA, present area, energy, and delay estimates for the NML PLA, compare the area of NML PLAs to other reprogrammable nanotechnologies, and analyze how architectural-level redundancy will affect performance and defect tolerance in NML PLAs. We will use results from this study to shape a concluding discussion about, which architectures appear to be most suitable for NML. Michael Crocker, Michael T. Niemier, Xiaobo Sharon Hu |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2011 | Performance and Energy Impact of Locally Controlled NML CircuitsabstractThis article quantitatively considers the performance of nanomagnetic logic circuits within the context of realistic drive circuitry. We also demonstrate how one of the five fundamental tenets of digital logic---preventing unwanted feedback---can be satisfied by realistic drive circuitry. More specifically, different types of multiphase clocks are investigated and compared. Initial projections suggest that even with drive circuitry overhead, nanomagnet logic can outperform subthreshold CMOS in terms of energy delay product---and paths to lower power exist. Aaron Dingler, Michael T. Niemier, Xiaobo Sharon Hu, Evan Lent |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2009 | Defects and faults in QCA-based PLAsabstractDefect tolerance will be critical in any system with nanoscale feature sizes. This article examines some fundamental aspects of defect tolerance for a reconfigurable system based on Quantum-dot Cellular Automata (QCA). We analyze a novel, QCA-based, Programmable Logic Array (PLA) structure, develop an implementation independent fault model, and discuss how expected defects and faults might affect yield. Within this context, we introduce techniques for mapping Boolean logic functions to a defective QCA-based PLA. Simulation results show that our new mapping techniques can achieve higher yields than existing techniques. Michael Crocker, Xiaobo Sharon Hu, Michael T. Niemier |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2008 | Bridging the gap between nanomagnetic devices and circuitsabstractThis paper looks at designing circuit elements that will be constructed with nanoscale magnets within the Quantum-dot Cellular Automata (QCA) computational paradigm. In magnetic QCA (MQCA) logical operations and dataflow are accomplished by manipulating the polarizations of nanoscale magnets. Wires and gates have already been experimentally demonstrated at room temperature. However, to realize more complex circuits - and eventually systems - more than just wires and gates in isolation are required. For example, gates must be inter-connected, signals must cross, etc. All structures must be controlled by the envisioned drive circuitry. In this paper, structures that will facilitate these circuit-level tasks are presented for the first time. Michael T. Niemier, Xiaobo Sharon Hu, Aaron Dingler, M. Tanvir Alam, Gary H. Bernstein, Wolfgang Porod |
ICCD | 1 |
| 2008 | Molecular QCA design with chemically reasonable constraintsabstractIn this article we examine the impacts of the fundamental constraints required for circuits and systems made from molecular Quantum-dot Cellular Automata (QCA) devices. Our design constraints are “chemically reasonable” in that we consider the characteristics and dimensions of devices and scaffoldings that have actually been fabricated. This work is a necessary first step for any work in QCA CAD, and can also help shape experiments in the physical sciences for emerging, nano-scale devices. Our work shows that QCA circuits, scaffoldings, substrates, and devices should all be considered simultaneously. Otherwise, there is a very real possibility that the devices and scaffoldings that are eventually manufactured will result in devices that only work in isolation. “Chemically reasonable” also means that expected manufacturing defects must be considered. In our simulations we introduce defects associated with self-assembled systems into various designs to begin to define manufacturing tolerances. This work is especially timely as experimentalists are beginning to work on merging experimental tracks that address devices and scaffolds—and the end result should facilitate correct logical operations. Michael Crocker, Michael T. Niemier, Xiaobo Sharon Hu, Marya Lieberman |
ACM J. Emerg. Technol. Comput. Syst. | 2 |
| 2007 | Fault Models and Yield Analysis for QCA-based PLAsabstractVarious implementations of the Quantum-dot Cellular Automata (QCA) device architecture may help many performance scaling trends continue as we approach the nano-scale. Experimental success has led to the evolution of a research track that looks at QCA-based design. The work presented in this paper follows that track and looks at implementation friendly, programmable QCA circuits. Specifically, we analyze a novel, QCA-based, Programmable Logic Array (PLA) structure, develop an implementation independent fault model, discuss how expected defects and faults might affect yield, and look at the design in the context of a magnetic implementation of QCA. Michael Crocker, Michael T. Niemier, Xiaobo Sharon Hu |
FPL | 2 |
| 2007 | Clocking structures and power analysis for nanomagnet-based logic devicesabstractLogical devices made from nano-scale magnets have many potential advantages - systems should be non-volatile, dense, low power, radiation hard, and could have a natural interface to MRAM. Initial work includes experimental demonstrations of logic gates and wires and theoretical studies that consider their power dissipation. This paper looks at power dissipation too, but also considers the circuitry needed to drive a computation. Initial results are very encouraging and indicate that clocked magnetic logic could - in the worst case - match equivalent low power CMOS circuits and - in the best-case - potentially provide more than 2 orders of magnitude improvement when one considers energy per operation. Michael T. Niemier, M. Alam, Xiaobo Sharon Hu, Gary H. Bernstein, Wolfgang Porod, M. Putney, J. DeAngelis |
ISLPED | 1 |
| 2007 | Approximating the Maximum Sharing Problem
Amitabh Chaudhary, Danny Ziyi Chen, Rudolf Fleischer, Xiaobo Sharon Hu, Jian Li 0015, Michael T. Niemier, Zhiyi Xie, Hong Zhu 0004 |
WADS | 6 |
| 2007 | Fabricatable Interconnect and Molecular QCA CircuitsabstractWhen exploring computing elements made from technologies other than complementary metal-oxide-semiconductor, it is imperative to investigate circuits and systems assuming realistic physical implementation constraints. This paper looks at molecular quantum-dot cellular automata (QCA) devices within this context. With molecular QCA, physical coplanar wire crossings may be very difficult to fabricate in the near to midterm. Here, we consider how this will affect interconnect. We introduce a novel technique to remove wire crossings in a given design in order to facilitate the self-assembly of real circuits - thus, providing meaningful and functional design targets for both physical and computer scientists. The proposed methodology eliminates all wire crossings with minimal logic gate/node duplications. Simulation results based on existing QCA circuits and other benchmarks are presented, and suggest that further investigation is needed. Amitabh Chaudhary, Danny Ziyi Chen, Xiaobo Sharon Hu, Michael T. Niemier, Ramprasad Ravichandran, Kevin Whitton |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 4 |
| 2006 | Using CAD to shape experiments in molecular QCAabstractThis paper examines how circuits and systems made from molecular QCA devices might function. Our design constraints are “chemically reasonable ” in that we consider the characteristics and dimensions of devices and scaffoldings (circuit boards to attach devices to) that have actually been fabricated (currently in isolation). We will show that not only is the work presented here a necessary first step for any work in QCA CAD, but also that by considering issues related to design can actually help shape experiments in the physical sciences for emerging, nano-scale devices. Our work shows that circuits, scaffoldings, substrates, and devices must all be considered simultaneously. Otherwise, there is a very real possibility that the devices and scaffoldings that are eventually manufactured will result in devices that only work in isolation. This work is especially timely as experimentalists are currently working to merge the different experimental tracks – i.e. to selectively place a QCA device. 1. Michael T. Niemier, Michael Crocker, Xiaobo Sharon Hu, Marya Lieberman |
ICCAD | 1 |
| 2005 | Partitioning and placement for buildable QCA circuitsabstractQuantum-dot Cellular Automata (QCA) is a novel computing mechanism that can represent binary information based on spatial distribution of electron charge configuration in chemical molecules. In this paper, we present partitioning and placement algorithms for a large-scale automatic QCA layout. The purpose of zone partitioning is to initially partition a given circuit such that a single clock potential modulates the interdot barriers in all of the QCA cells within each zone. We then place these zones during our placement step. We identify several objectives and constraints that will enhance the buildability of QCA circuits and use them in our optimization process. The results are intended to define what is computationally interesting and could actually be built within a set of predefined constraints. Ramprasad Ravichandran, Michael T. Niemier, Sung Kyu Lim |
ASP-DAC | 2 |
| 2005 | Eliminating wire crossings for molecular quantum-dot cellular automata implementationabstractWhen exploring computing elements made from technologies other than CMOS, it is imperative to investigate the effects of physical implementation constraints. This paper focuses on molecular quantum-dot cellular automata circuits. For these circuits, it is very difficult for chemists to fabricate wire crossings (at least in the near future). A novel technique is introduced to remove wire crossings in a given circuit to facilitate the self assembly of real circuits - thus providing meaningful and functional design targets for both physical and computer scientists. The technique eliminates all wire crossings with minimal logic gate/node duplications. Experimental results based on existing QCA circuits and other benchmarks are quite encouraging, and suggest that further investigation is needed. Amitabh Chaudhary, Danny Ziyi Chen, Kevin Whitton, Michael T. Niemier, Ramprasad Ravichandran |
ICCAD | 4 |
| 2005 | Automatic cell placement for quantum-dot cellular automata
Ramprasad Ravichandran, Sung Kyu Lim, Michael T. Niemier |
Integr. | 3 |
| 2005 | Partitioning and placement for buildable QCA circuitsabstractQuantum-dot Cellular Automata (QCA) is a novel computing mechanism that can represent binary information based on spatial distribution of an electron charge configuration in chemical molecules. In this article, we present the first partitioning and placement algorithm for automatic QCA layout. We identify several objectives and constraints that will enhance the buildability of QCA circuits. The results are intended to: (1) define what is computationally interesting and could actually be built within a set of predefined constraints, (2) project what designs will be possible as additional constructs become realizable, and (3) provide a vehicle that we can use to compare QCA systems to silicon-based systems. Sung Kyu Lim, Ramprasad Ravichandran, Michael T. Niemier |
ACM J. Emerg. Technol. Comput. Syst. | 3 |
| 2004 | Quantum-Dot Cellular Automata (QCA) circuit partitioning: problem modeling and solutionsabstractThis paper presents the Quantum-Dot Cellular Automata (QCA) physical design problem, in the context of the VLSI physical design problem. The problem is divided into three subproblems: partitioning, placement, and routing of QCA circuits. This paper presents an ILP formulation and heuristic solution to the partitioning problem, and compares the two sets of results. Additionally, we compare a human-generated circuit to the ILP and Heuristic solutions. The results demonstrate that the heuristic is a practical method of reducing partitioning run time while providing a result that is close to the optimal for a given circuit. Dominic A. Antonelli, Danny Ziyi Chen, Timothy J. Dysart, Xiaobo Sharon Hu, Andrew B. Kahng, Peter M. Kogge, Richard C. Murphy, Michael T. Niemier |
DAC | 8 |
| 2004 | Automatic cell placement for quantum-dot cellular automataabstractQuantum-dot Cellular Automata (QCA) is a novel computing mechanism that can represent binary information based on spatial distribution of electron charge configuration in chemical molecules. It has the potential to allow for circuits and systems with functional densities that are better than end of the roadmap CMOS, but also imposes new constraints on system designers. In this paper we develop the first cell-level placement of QCA circuits, where the given circuit is assumed to be partitioned into 4-phase asynchronous QCA timing zones. We formulate the QCA cell placement in each timing zone as a unidirectional geometric embedding of k-layered bipartite graphs. We then present an analytical and a stochastic solution for minimizing the wire crossings and wire length in these placement solutions. Ramprasad Ravichandran, Nihal Ladiwala, Jean Nguyen, Michael T. Niemier, Sung Kyu Lim |
ACM Great Lakes Symposium on VLSI | 4 |
| 2004 | Using Circuits and Systems-Level Research to Drive NanotechnologyabstractThis paper details nano-scale devices being researched by physical scientists to build computational systems. It also reviews some existing system design work that uses the devices to be discussed. It concludes with a discussion of how the authors believe system-level research can best be used to positively affect actual device development. This work has led to a more thorough design methodology that address whether or not computationally interesting and buildable circuits are possible with the quantum-dot cellular automata (QCA), while also providing significant wins over end-of-the-roadmap CMOS. Michael T. Niemier, Ramprasad Ravichandran, Peter M. Kogge |
ICCD | 1 |
| 2001 | Exploring and exploiting wire-level pipelining in emerging technologiesabstractPipelining is a technique that has long since been considered fundamental by computer architects. However, the world of nanoelectronics is pushing the idea of pipelining to new and lower levels — particularly the device level. How this affects circuits and the relationship between their timing, architecture, and design will be studied in the context of an inherently self-latching nanotechnology termed Quantum Cellular Automata (QCA). Results indicate that this nanotechnology offers the potential for “free” multi-threading and “processing-in-wire”. All of this could be accomplished in a technology that could be almost three orders of magnitude denser than an equivalent design fabricated in a process at the end of the CMOS curve. Michael T. Niemier, Peter M. Kogge |
ISCA | 1 |
| 2000 | A design of and design tools for a novel quantum dot based microprocessorabstractDespite the seemingly endless upw ards spiral of modern VLSI technology, many experts are predicting a hard w all for CMOS in about a decade. Given this, researc hers con tin ue to look at alternative technologies, one of which is based on quan tumdots, called quan tumcellular automata (QCA). While the first such devices have been fabricated, little is kno wn about how to design complete systems of them. This paper summarizes one of the first such studies, namely an attempt to design a complete, albeit simple, CPU in the technology. T o design a theoretical QCA microprocessor, two things must be accomplished. First a device model of the processor must be constructed (i.e. the schematic itself). Second, methods for sim ulatingand testing QCA designs m ust be developed. This paper summarizes the beginnings of a simple QCA microprocessor (namely, its dataflow) and a QCA design and simulation tool. Michael T. Niemier, Michael J. Kontz, Peter M. Kogge |
DAC | 1 |
| 1999 | Logic in Wire: Using Quantum Dots to Implement a MicroprocessorabstractDespite the seemingly endless upwards spiral of modern VLSI technology many experts are predicting a hard wall for CMOS in about a decade. Given this, researchers continue to look at alternative technologies, one of which is based on quantum dots, called quantum cellular automata. While the first such devices have been fabricated, little is known about how to design complete systems. This paper summarizes one of the first such studies, namely an attempt to design a complete, albeit simple, CPU in the technology. The projections are striking: a projected 10 to 1 increase in circuit density when compared to a CMOS equivalent, but a design approach which is radically different from conventional "logic" design, especially in timing considerations. Michael T. Niemier, Peter M. Kogge |
Great Lakes Symposium on VLSI | 1 |