VLDB 2026 Research / reviewers in the wild / expert
Dayane Reis
dblp:223/9620 · also Dayane Alfenas Reis
· DBLP profile ↗
28ranked-venue papers
9as first author
18since 2021 · last 2026
0000-0002-8571-1308ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 25 · 9 first-author · 16 since 2021Software engineering, systems software and programming languages · 5 · 2 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Security and privacy · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Late Breaking Results: SP-HD: Stochastic Projection-Based HyperDimensional Architecture for Near-Sensor Image ClassificationabstractThis paper presents SP-HD, a near-sensor image classification architecture that combines stochastic computing (SC) and hyperdimensional computing (HDC) to enable energy-efficient and compact embedded intelligence. The proposed approach introduces a stochastic projection mechanism that converts input features into bitstreams, enabling bipolar multiplications to be performed with simple logic and in-memory accumulation, thereby eliminating costly multipliers and level hypervectors. A mixed-signal ReRAM-based implementation further reduces data movement by performing projection and accumulation directly within the memory fabric, while binary-weight classification minimizes circuit complexity. SP-HD achieves competitive accuracy across multiple image datasets and delivers 3μJ energy per inference with a compact 2.56mm2hardware footprint, significantly outperforming prior ReRAM compute-in-memory accelerators in both energy and area efficiency. Ahmed Mamdouh, Sabrina Hassan Moon, Abu Kaisar Mohammad Masum, Emilien Meyer, Sercan Aygün, Dayane Reis |
DATE | 6 |
| 2026 | Increasing the Efficiency of Associative Processor Architectures via CMOS-Compatible HybridizationabstractWe present a hybrid, general-purpose, associative processing-in-memory architecture that combines the energy and area advantages of a primary FeFET-based CAM array with the write performance and endurance of a much smaller CMOS-based sidekick. The hybrid nature of the architecture is transparent to the programmer, who uses a RISC-V ISA with standard RVV vector extensions. Detailed SPICE- and system-level simulations show our hybrid design dramatically curbs the endurance disadvantages of a pure FeFET design and delivers, on average, 30% and 11% area and energy savings over a purely CMOS implementation, respectively, at a performance loss of barely 1% over pure CMOS. Socrates S. Wong, Cecilio C. Tamarit, Mohammad Mehdi Sharifi, Zephan M. Enciso, Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu, José F. Martínez |
DATE | 5 |
| 2026 | SATurn: A Low-Power FeFET Crossbar Architecture for Solving Boolean Satisfiability ProblemsabstractBoolean satisfiability (SAT) continues to serve as the fundamental NP-Complete problem of computer science, where many practical problems in cryptography, circuit design, and verification have been translated into large SAT problems. Many hardware approaches rely on clause evaluation, which can become computationally infeasible on traditional von Neumann architectures. Crossbar arrays with emerging nonvolatile memory (NVM) devices provide a fast solution for vector-matrix multiplication (VMM), and clause evaluation can be expressed as a VMM. Leveraging the power of Computing-in-Memory (CiM) we introduce SATurn, a ferroelectric field-effect transistor (FeFET)-based architecture for performing stochastic local search (SLS) with the WalkSAT-XNF heuristic on SAT problems. We compare our design with two other CiM SLS SAT solver architectures, achieving energy consumption improvements of up to 5.72 × for clause evaluation, 6.52 × for heuristic calculation, and 20.63 × for the make/break value computations under the same latency constraints. John Taylor Maurer, Ahmed Mamdouh Mohamed Ahmed, Parsa Khorrami, Dayane Reis |
ACM Great Lakes Symposium on VLSI | 4 |
| 2026 | XL-HD: Extended Learning in Hyperdimensional Computing via Deterministic Projections for In-Memory AcceleratorsabstractHyperdimensional computing (HDC) is a promising approach for energy-efficient edge machine learning (ML), where low latency, low power, and tight memory budgets are essential. However, traditional HDC relies on symbolic binding and pseudo-random high-dimensional vectors, which require large dimensionality and heuristic updates to reach competitive accuracy, limiting deployment on edge hardware. We introduce XL-HD, a deterministic, projection-based, fully learnable HDC framework tailored for in-memory acceleration within edge computing systems. The method uses a fixed Sobol sequence to project binary inputs, extending learning beyond conventional HDC. During training, class prototypes are optimized in real-valued space and later binarized, enabling an entirely binary dot-product inference pipeline ideal for IMC hardware such as ReRAM crossbars. XL-HD achieves competitive accuracy on MNIST, UCIHAR, and ISOLET while maintaining a compact IMC-based inference engine with 0.395 mm2 area and only 0.40 μJ per single-cycle inference. Sabrina Hassan Moon, Abu Kaisar Mohammad Masum, Sercan Aygün, Dayane Reis |
ISLPED | 4 |
| 2026 | Enhancing biologically inspired hierarchical temporal memory with hardware-accelerated reflex memory
Pavia Bera, Sabrina Hassan Moon, Jennifer Adorno, Dayane Reis, Sanjukta Bhanja |
Neurocomputing | 4 |
| 2026 | PPIMCE: In-Memory Computing Fabric for Privacy Preserving Computing
Jianqiao Mo, Dayane Reis, Jonathan Takeshita, Taeho Jung, Brandon Reagen, Michael T. Niemier, Xiaobo Sharon Hu |
J. Comput. Sci. Technol. | 3 |
| 2025 | Late Breaking Results: On-the-Fly Hadamard Hypervector Processing for Efficient Hyperdimensional ComputingabstractInspired by the human brain, Hyperdimensional Computing (HDC) processes information efficiently by operating in high-dimensional space using hypervectors. While previous works focus on optimizing pregenerated hypervectors in software, this study introduces a novel on-the-fly vector generation method in hardware with $O(1)$ complexity, compared to the $O(N)$ iterative search used in conventional approaches to find the best orthogonal hypervectors. Our approach leverages Hadamard binary coefficients and unary computing to simplify encoding into addition-only operations after the generation stage in ASIC, implemented using inmemory computing. The proposed design significantly improves accuracy and computational efficiency across multiple benchmark datasets. Abu Kaisar Mohammad Masum, Mehran Shoushtari Moghadam, Sabrina Hassan Moon, Ahmed Mamdouh Mohamed Ahmed, M. Hassan Najafi, Dayane Reis, Sercan Aygün |
DAC | 6 |
| 2025 | ReX-HD: A Deterministic ReRAM-Based Hyperdimensional Computing Framework for Edge Computing
Sabrina Hassan Moon, Ahmed Mamdouh, Abu Kaisar Mohammad Masum, Sercan Aygün, Dayane Reis |
ACM Great Lakes Symposium on VLSI | 5 |
| 2025 | Shared-PIM: Enabling Concurrent Computation and Data Flow for Faster Processing-in-DRAMabstractProcessing-in-Memory (PIM) enhances memory with computational capabilities, potentially solving energy and latency issues associated with data transfer between memory and processors. However, managing concurrent computation and data flow within the PIM architecture incurs significant latency and energy penalty for applications. This paper introduces Shared-PIM, an architecture for in-DRAM PIM that strategically allocates rows in memory banks, bolstered by memory peripherals, for concurrent processing and data movement. Shared-PIM enables simultaneous computation and data transfer within a memory bank. When compared to LISA, a state-of-the-art architecture that facilitates data transfers for in-DRAM PIM, Shared-PIM reduces data movement latency and energy by 5× and 1.2×, respectively. Furthermore, when integrated to a state-of-the-art (SOTA) in-DRAM PIM architecture (pLUTo), Shared-PIM achieves 1.4× faster addition and multiplication, and thereby improves the performance of matrix multiplication (MM) tasks by 40%, polynomial multiplication (PMM) by 44%, and numeric number transfer (NTT) tasks by 31%. Moreover, for graph processing tasks like Breadth-First Search (BFS) and Depth-First Search (DFS), Shared-PIM achieves a 29% improvement in speed, all with an area overhead of just 7.16% compared to the baseline pLUTo. Ahmed Mamdouh, Michael T. Niemier, Xiaobo Sharon Hu, Dayane Reis |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 5 |
| 2024 | AFeCAM: An Energy Efficient Analog 1FeFET Content Addressable MemoryabstractContent Addressable Memories (CAMs) have the ability to perform parallel searches, significantly enhancing the computational efficiency of Computing-in-Memory (CiM) architectures. CAMs can be employed in various areas, including DNA sequence analysis, IP routing, etc. Meanwhile, Ferroelectric Field Effect Transistors (FeFETs) offer a highly efficient solution for CAM implementation due to their high Ion/Ioff ratio, voltage-driven write mechanism, and non-volatility. We propose a 1FeFET analog CAM design (AFeCAM), which minimizes the area footprint compared to previous FeFET-based CAMs, making it possible to conduct searches with minimal search energy in data-intensive applications. Our AFeCAM design is ultra-compact, scalable, and uses a voltage comparator sense amplifier to detect matches and mismatches in an energy-efficient manner. Sabrina Hassan Moon, Dayane Reis |
ACM Great Lakes Symposium on VLSI | 2 |
| 2024 | Accelerating Finite-Field and Torus Fully Homomorphic Encryption via Compute-Enabled (S)RAMabstractFully Homomorphic Encryption (FHE) allows outsourced computation on clients’ encrypted data while preserving data privacy. FHE’s high computational intensity incurs high overhead from data transfer with hardware such as CPU, GPU, and FPGA, due to the inherent separation between computing and data. To overcome this limitation, Compute-Enabled RAM (CE-RAM) has been explored; however, prior work using CE-RAM to accelerate FHE only explores a simple implementation of a finite-field FHE scheme and did not explore algorithmic optimizations.In this paper, we investigate CE-RAM acceleration FHE more deeply, implementing both the finite-field B/FV and torus-based TFHE cryptosystems in CE-RAM with common FHE optimizations. This is the first work to explore using CE-RAM to accelerate TFHE. For B/FV, we explore parameter-specific algorithmic optimizations specifically designed for CE-RAM friendliness. We evaluate our implementation as compared to prior work in CE-RAM FHE acceleration and other hardware acceleration strategies. We demonstrate speedups of up to 784x for B/FV homomorphic multiplication and 38x for TFHE bootstrapping as compared to CPU implementations. We also discuss the overhead of CE-RAM for FHE on energy and area consumption, showing comparable or improved performance as compared to other work or hypothetical near-memory accelerators. Jonathan Takeshita, Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu, Taeho Jung |
IEEE Trans. Computers | 2 |
| 2024 | A Computing-in-Memory-Based One-Class Hyperdimensional Computing Model for Outlier DetectionabstractIn this work, we presentODHD, an algorithm for outlier detection based on hyperdimensional computing (HDC), a non-classical learning paradigm. Along with the HDC-based algorithm, we proposeIM-ODHD, a computing-in-memory (CiM) implementation based on hardware/software (HW/SW) codesign for improved latency and energy efficiency. The training and testing phases ofODHDmay be performed with conventional CPU/GPU hardware or ourIM-ODHD, SRAM-based CiM architecture using the proposed HW/SW codesign techniques. We evaluate the performance ofODHDon six datasets from different application domains using three metrics, namely accuracy, F1 score, and ROC-AUC, and compare it with multiple baseline methods such as OCSVM, isolation forest, and autoencoder. The experimental results indicate thatODHDoutperforms all the baseline methods in terms of these three metrics on every dataset for both CPU/GPU and CiM implementations. Furthermore, we perform an extensive design space exploration to demonstrate the tradeoff between delay, energy efficiency, and performance ofODHD. We demonstrate that the HW/SW codesign implementation of the outlier detection onIM-ODHDis able to outperform the GPU-based implementation ofODHDby at least 331.5×/889× in terms of training/testing latency (and on average 14.0×/36.9× in terms of training/testing energy consumption). Sabrina Hassan Moon, Xiaobo Sharon Hu, Xun Jiao 0002, Dayane Reis |
IEEE Trans. Computers | 5 |
| 2023 | In-Memory Computing Accelerators for Emerging Learning ParadigmsabstractOver the past decades, emerging, data-driven machine learning (ML) paradigms have increased in popularity, and revolutionized many application domains. To date, a substantial effort has been devoted to devising mechanisms for facilitating the deployment and near ubiquitous use of these memory intensive ML models. This review paper presents the use of in-memory computing (IMC) accelerators for emerging ML paradigms from a bottom-up perspective through the choice of devices, the design of circuits/architectures, to the application-level results. Dayane Reis, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu |
ASP-DAC | 1 |
| 2023 | Invited Paper: Algorithm/Hardware Co-Design for Few-Shot Learning at the EdgeabstractOn-device learning is essential to achieve intelligence at the edge, where it is desirable to learn from few samples or even just a single sample. Memory-augmented neural networks (MANNs), which augment neural networks with an attentional memory, can draw on already learnt knowledge patterns and adapt to new but similar tasks. Implementing MANNs on conventional architectures can require a significant amount of costly data transfer, thereby limiting the practical use of MANNs at the edge. In this paper, we introduce algorithm/hardware co-design solutions which exploit compact designs of content addressable memories (CAMs) based on emerging non-volatile memories (e.g., FeFETs) to implement energy-efficient MANN accelerators. The design space of MANN accelerators is systematically analyzed by considering different circuit, architecture, and algorithm options. We further discuss how hyper-dimensional representations of data can be combined with MANNs to overcome the negative effect of device/circuit variabilities on learning quality, thus achieving not only energy-efficient but also accuracy-competitive on-device learning at the edge. We also investigate modeling of device-to-device (D2D) variation in FeFETs using the write-with-verify approach and detail its impact on the energy, delay, and accuracy of the MANN application. Ann Franchesca Laguna, Mohammad Mehdi Sharifi, Dayane Reis, Liu Liu 0023, Andrew Hennessee, Clayton O'Dell, Ian O'Connor, Michael T. Niemier, Xiaobo Sharon Hu |
ICCAD | 3 |
| 2022 | iMARS: an in-memory-computing architecture for recommendation systemsabstractRecommendation systems (RecSys) suggest items to users by predicting their preferences based on historical data. Typical RecSys handle large embedding tables and many embedding table related operations. The memory size and bandwidth of the conventional computer architecture restrict the performance of RecSys. This work proposes an in-memory-computing (IMC) architecture (iMARS) for accelerating the filtering and ranking stages of deep neural network-based RecSys. iMARS leverages IMC-friendly embedding tables implemented inside a ferroelectric FET based IMC fabric. Circuit-level and system-level evaluation show that iMARS achieves 16.8x (713x) end-to-end latency (energy) improvement compared to the GPU counterpart for the MovieLens dataset. Mengyuan Li 0001, Ann Franchesca Laguna, Dayane Reis, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu |
DAC | 3 |
| 2022 | IMCRYPTO: An In-Memory Computing Fabric for AES Encryption and DecryptionabstractThis article proposes IMCRYPTO, an in-memory computing (IMC) fabric for accelerating advanced encryption standard (AES) encryption and decryption. IMCRYPTO employs a unified structure to implement encryption and decryption in a single-hardware architecture with combined (Inv)SubBytes and (Inv)MixColumns steps. Because of this step combination and the high parallelism achieved by multiple units of random access memory (RAM) and random access/content addressable memory (RA/CAM) arrays, IMCRYPTO achieves high-throughput encryption and decryption without sacrificing area and power consumption. In addition, due to the integration of an RISC-V core, IMCRYPTO offers programmability and flexibility. IMCRYPTO improves the throughput per area by a minimum (maximum) of$3.3\times $($223.1\times $) compared to previous ASICs/IMC architectures for AES-128 encryption. Projections show added benefit from emerging technologies of up to$5.3\times $to the area–delay–power product of IMCRYPTO. Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2021 | Attention-in-Memory for Few-Shot Learning with Configurable Ferroelectric FET ArraysabstractAttention-in-Memory (AiM), a computing-in-memory (CiM) design, is introduced to implement the attentional layer of Memory Augmented Neural Networks (MANNs). AiM consists of a memory array based on Ferroelectric FETs (FeFET) along with CMOS peripheral circuits implementing configurable functionalities, i.e., it can be dynamically changed from a ternary content-addressable memory (TCAM) to a general-purpose (GP) CiM. When compared to state-of-the art accelerators, AiM achieves comparable end-to-end speed-up and energy for MANNs, with better accuracy (95.14% v.s. 92.21%, and 95.14% v.s. 91.98%) at iso-memory size, for a 5-way 5-shot inference task with the Omniglot dataset. Dayane Reis, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu |
ASP-DAC | 1 |
| 2021 | Exploiting FeFETs via Cross-Layer Design from In-memory Computing Circuits to Meta-Learning ApplicationsabstractA ferroelectric FET (FeFET), made by integrating a ferroelectric material layer in the gate stack of a MOSFET, is a device that can behave as both a transistor and a non-volatile storage element. This unique property of FeFETs enables area efficient and low-power merged logic and memory functionality, desirable for many data analytic and machine learning applications. To best exploit this unique feature of FeFETs, cross-layer design practices spanning from circuits and architectures to algorithms and applications is needed. The paper presents FeFET-based circuits and architectures that offer, either independently or in a configurable fashion, content addressable memory (TCAM) and general-purpose compute-in-memory (GP-CiM) functionalities. These in-memory computing modules bring new opportunities to accelerating data-intensive applications. We discuss the use of these FeFET based in-memory computing fabrics in meta-learning applications, specifically as attentional memory. System-level task mapping and end-to-end evaluation will be discussed. Dayane Reis, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 1 |
| 2020 | Emerging Neural Workloads and Their Impact on HardwareabstractWe consider existing and emerging neural workloads, and what hardware accelerators might be best suited for said workloads. We begin with a discussion of analog crossbar arrays, which are known to be well-suited for matrix-vector multiplication operations that are commonplace in existing neural network models such as convolutional neural networks (CNNs). We highlight candidate crosspoint devices, what device and materials challenges must be overcome for a given device to be employed in a crossbar array for a computationally interesting neural workload, and how circuit and algorithmic optimizations may be employed to mitigate undesirable characteristics from devices/materials. We then discuss two emerging neural workloads. We first consider machine learning models for one- and few-shot learning tasks (i.e., where a network can be trained with just one or a few, representative examples of a given class). Notably crossbar-based architectures can be used to accelerate said models. Hardware solutions based on content addressable memory arrays will also be discussed. We then consider machine learning models for recommendation systems. Recommendation models, an emerging class of machine learning models, employ distinct neural network architectures that operate of continuous and categorical input features which make hardware acceleration challenging. We will discuss the open research challenges and opportunities within this space. David Brooks 0001, Martin M. Frank, Tayfun Gokmen, Udit Gupta 0001, Xiaobo Sharon Hu, Shubham Jain 0004, Ann Franchesca Laguna, Michael T. Niemier, Ian O'Connor, Anand Raghunathan, Ashish Ranjan 0001, Dayane Reis, Jacob R. Stevens, Carole-Jean Wu, Xunzhao Yin |
DATE | 12 |
| 2020 | A Fast and Energy Efficient Computing-in-Memory Architecture for Few-Shot Learning ApplicationsabstractAmong few-shot learning methods, prototypical networks (PNs) are one of the most popular approaches due to their excellent classification accuracies and network simplicity. Test examples are classified based on their distances from class prototypes. Despite the application-level advantages of PNs, the latency of transferring data from memory to compute units is much higher than the PN computation time. Thus, PNs performance is limited by memory bandwidth. Computing-in-memory addresses this bandwidth-bottleneck problem by bringing a subset of compute units closer to memory. In this work, we propose a CiM-PN framework that enables the computation of distance metrics and prototypes inside the memory. CiM-PN replaces the computationally intensive Euclidean distance metric by the CiM-friendly Manhattan distance metric. Additionally, prototypes are computed using an in-memory mean operation realized by accumulation and division by powers of two, which enables few-shot learning implementations where "shots" are powers of two. The CiM-PN hardware uses CMOS memory cells, as well as CMOS peripherals such as customized sense amplifiers, carry-look-ahead adders, in-place copy buffers and a logarithmic shifter. Compared with a GPU implementation, a CMOS-based CiM-PN achieves speedups of 2808x/111x and energy savings of 2372x/5170x at iso-accuracy for the prototype and nearest-neighbor computation, respectively, and over 2x end-to-end speedup and energy improvements. We also gain 3-14% accuracy improvement when compared to existing non-GPU hardware approaches due to the floating-point CiM operations. Dayane Reis, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu |
DATE | 1 |
| 2020 | Modeling and Benchmarking Computing-in-Memory for Design Space ExplorationabstractThe bottleneck between the limited memory bandwidth and high speed processing demands is the main cause of problems associated with high volume of data transfers in data-intensive applications. As a possible remedy to these issues, computing-in-memory (CiM) enables a subset of logic and arithmetic operations to be performed where the data resides, i.e., inside the memory. Various CiM designs have been proposed to date, based on different technologies. Given the variety of options available, picking the right design option for a system/application can be a complex task. When choosing a CiM design, it is important to establish evaluation conditions that are as uniform as possible to make a fair choice between available design options. In this paper, we describe a methodology for an uniform benchmarking of CiM designs. Our approach evaluates devices/circuits, arrays and the overall impact of CiM to a system with a framework based on Eva-CiM. As a case study, we analyze the array-level performance of 7 recent CiM designs implemented with SRAM, DRAM, FeFET-RAM, STT-MRAM, SOT-MRAM, and RRAM. After we identify that the FeFET-RAM-based design shows promising energy and delay savings at the array level, we carry out a system level evaluation showing that FeFET-RAM-based CiM outperforms a CMOS SRAM CiM baseline by an average of 60% across a set of 17 benchmarks (with respect to energy savings). Regarding speedups, both technologies offer virtually the same benefit of about 1.5X when compared to a situation where processing does not happen in memory. Dayane Reis, Shaahin Angizi, Xunzhao Yin, Deliang Fan, Michael T. Niemier, Cheng Zhuo, Xiaobo Sharon Hu |
ACM Great Lakes Symposium on VLSI | 1 |
| 2020 | Algorithmic Acceleration of B/FV-Like Somewhat Homomorphic Encryption for Compute-Enabled RAM
Jonathan Takeshita, Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu, Taeho Jung |
SAC | 2 |
| 2020 | Eva-CiM: A System-Level Performance and Energy Evaluation Framework for Computing-in-Memory ArchitecturesabstractComputing-in-memory (CiM) architectures aim to reduce costly data transfers by performing arithmetic and logic operations in memory and hence relieve the pressure due to the memory wall. However, determining whether a given workload can really benefit from CiM, which memory hierarchy and what device technology should be adopted by a CiM architecture requires in-depth study that is not only time consuming but also demands significant expertise in architectures and compilers. This article presents an energy and performance evaluation framework, Eva-CiM, for systems based on CiM architectures. Eva-CiM encompasses a multilevel (from device to architecture) comprehensive tool chain that leverages existing modeling and simulation tools, such as GEM5, McPAT, and DESTINY. To support high-confidence prediction, rapid design space exploration and ease of use, Eva-CiM introduces several novel modeling/analysis approaches including models for capturing memory access and dependency-aware ISA traces, and for quantifying interactions between the host CPU and the CiM module. Eva-CiM can readily produce energy and performance estimates of the entire system for a given program, a processor architecture, and the CiM array and technology specifications. Eva-CiM is validated by comparing with DESTINY. Eva-CiM enables analyses including the system-level impact of CiM-supported accesses, whether a program is CiM-favorable as well as the pros and cons of increased memory size for CiM. Eva-CiM also facilitates exploration of different design configurations and technologies. Dayane Reis, Xiaobo Sharon Hu, Cheng Zhuo |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2020 | Computing-in-Memory for Performance and Energy-Efficient Homomorphic EncryptionabstractHomomorphic encryption (HE) allows direct computations on encrypted data. Despite numerous research efforts, the practicality of HE schemes remains to be demonstrated. In this regard, the enormous size of ciphertexts involved in HE computations degrades computational efficiency. Near-memory processing (NMP) and computing-in-memory (CiM)—paradigms where computation is done within the memory boundaries—represent architectural solutions for reducing latency and energy associated with data transfers in data-intensive applications, such as HE. This article introduces CiM-HE, a CiM architecture that can support operations for the Brakerski/Fan–Vercauteren (B/FV) scheme, a somewhat HE scheme for general computation. CiM-HE hardware consists of customized peripherals, such as sense amplifiers, adders, bit shifters, and sequencing circuits. The peripherals are based on CMOS technology and could support computations with memory cells of different technologies. Circuit-level simulations are used to evaluate our CiM-HE framework assuming a 6T-SRAM memory. We compare our CiM-HE implementation against: 1) two optimized CPU HE implementations and 2) a field-programmable gate array (FPGA)-based HE accelerator implementation. Compared with a CPU solution, CiM-HE obtains speedups between$4.6\times $and$9.1\times $and energy savings between$266.4\times $and$532.8\times $for homomorphic multiplications (the most expensive HE operation). Also, a set of four end-to-end tasks, i.e., mean, variance, linear regression, and inference, are up to$1.1\times $,$7.7\times $,$7.1\times $, and$7.5\times $faster (and$301.1\times $,$404.6\times $,$532.3\times $, and$532.8\times $more energy efficient). Compared with CPU-based HE in previous work, CiM-HE obtains$14.3\times $speedup and$> 2600\times $energy savings. Finally, our design offers$2.2\times $speedup with$88.1\times $energy savings compared with a state-of-the-art FPGA-based accelerator. Dayane Reis, Jonathan Takeshita, Taeho Jung, Michael T. Niemier, Xiaobo Sharon Hu |
IEEE Trans. Very Large Scale Integr. Syst. | 1 |
| 2019 | Ferroelectric FET Based In-Memory Computing for Few-Shot LearningabstractAs CMOS technology advances, the performance gap between the CPU and main memory has not improved. Furthermore, the hardware deployed for Internet of Things (IoT) applications need to process ever growing volumes of data, which can further exacerbate the "memory wall". Computing-in-memory (CiM) architectures, where logic and arithmetic operations are performed in memory, can significantly reduce energy and latency overheads associated with data transfer, and potentially alleviate processor-memory bottlenecks. In this paper, we consider the utility of ternary content addressable memory (TCAM) arrays and CiM arrays based on ferroelectric field effect transistors (FeFETs) to support emerging machine learning models that can learn new classes of data with significantly less training overhead - highly desirable in IoT applications. Architecturally, we use TCAM and CiM arrays to implement the external memory module in a memory enhanced neural network (MENN) - which can be used to minimize catastrophic forgetting - a major problem in applications such as lifelong and few-shot learning. As a representative example, we achieve 95.14% accuracy for a few-shot learning task with the Omniglot data set by using a combined L∞ infinity and L1 distance metric computed via a TCAM-CiM cascaded architecture (as opposed to 99.06% accuracy assuming a GPU backed by DRAM). While there is a slight drop in accuracy, the TCAM-CiM approach is 4.34X faster and 4.18X more energy efficient than a CMOS implementation for the same task. The ability of an FeFET to serve as both a compact logic and storage element helps to enable dense CiM and TCAM structures that drive the aforementioned improvements to application-level figures of merit (FOMs). Ann Franchesca Laguna, Xunzhao Yin, Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu |
ACM Great Lakes Symposium on VLSI | 3 |
| 2019 | The Impact of Emerging Technologies on Architectures and System-level Management: Invited PaperabstractThe goal of this work is to introduce and discuss different kinds of emerging technologies for logic circuitry and memory with respect to the key question of how they will impact future system-on-chip architectures and system-level management techniques. It is obvious that emerging technologies should have an impact there in order to fully exploit their technological advantages but also in order to deal with any disadvantages they might come with. In this special session paper, three promising emerging technologies are presented: (i) Negative Capacitance Field-Effect Transistor (NCFET) as a new CMOS technology with advantages primarily for low-power design, (ii) Ferroelectric FET (FeFET) as a non-volatile, area-efficient and low-power combined logic and memory as well as (iii) a Phase-Change Memory (PCM) and Resistive RAM (ReRAM) offering a large potential for tackling the memory wall problem in the von Neumann architecture. Our analysis demonstrates that not only new computing paradigms are promoted by these new technologies, it will also be seen that the trade-offs between the classical design parameters of low power, performance etc. will shift and hence emerging technologies will offer new Pareto points in the design space of future on-chip architectures. In that context, this work is unique as it bridges the gap between the technology side and system/architecture-level side to draw a vision of new technologies and their impact on architectures and system-level management. Jörg Henkel, Hussam Amrouch, Martin Rapp, Sami Salamin, Dayane Reis, Xunzhao Yin, Michael T. Niemier, Cheng Zhuo, Xiaobo Sharon Hu, Hsiang-Yun Cheng, Chia-Lin Yang |
ICCAD | 5 |
| 2018 | Computing in memory with FeFETsabstractData transfer between a processor and memory frequently represents a bottleneck with respect to improving application-level performance. Computing in memory (CiM), where logic and arithmetic operations are performed in memory, could significantly reduce both energy consumption and computational overheads associated with data transfer. Compact, low-power, and fast CiM designs could ultimately lead to improved application-level performance. This paper introduces a CiM architecture based on ferroelectric field effect transistors (FeFETs). The CiM design can serve as a general purpose, random access memory (RAM), and can also perform Boolean operations ((N)AND, (N)OR, X(N)OR, INV) as well as addition (ADD) between words in memory. Unlike existing CiM designs based on other emerging technologies, FeFET-CiM accomplishes the aforementioned operations via a single current reference in the sense amplifier, which leads to more compact designs and lower power. Furthermore, the high Ion/Ioff ratio of FeFETs enables an inexpensive voltage-based sense scheme. Simulation-based case studies suggest that our FeFET-CiM can achieve speed-ups (and energy reduction) of ~119X (~1.6X) and ~1.97X (~1.5X) over ReRAM and STT-RAM CiM designs with respect to in-memory addition of 32-bit words. Furthermore, our approach offers an average speedup of ~2.5X and energy reduction of ~1.7X when compared to a conventional (not in-memory) approach across a wide range of benchmarks. Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu |
ISLPED | 1 |
| 2016 | A Methodology for Standard Cell Design for QCAabstractQCA (Quantum-Dot Cellular Automata) is an emerging nanotechnology that has the potential to replace current CMOS technologies. QCA permits extremely low power consumption, since its working principle is not based on electric current flow but on Coulomb interaction. The development of Electronic Design Automation (EDA) tools and flows is an essential step towards the applicability of QCA for integrated designs. However, the scarce number of works in this field highlights that there is plenty of room for the development of new EDA methodologies for emerging nanotechnologies. Standard cells play an important role in this context, since the development of routing and placement algorithms are strongly related to their existence. This work presents a methodology for standard cell design for QCA as well as the exemplary QCA cell library ONE, which is based on the recently proposed USE (Universal, Scalar and Efficient) clocking scheme. Two representative case studies indicate the feasibility of the approach. Dayane Reis, Caio Araujo T. Campos, Thiago Rodrigues B. S. Soares, Omar P. Vilela Neto, Frank Sill |
ISCAS | 1 |