Ann Franchesca Laguna

dblp:150/9555 · also Ann Franchesca B. Laguna · DBLP profile ↗
← Back
23ranked-venue papers
5as first author
18since 2021 · last 2025
0000-0001-8267-1040ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 20 · 5 first-author · 15 since 2021Software engineering, systems software and programming languages · 7 · 2 first-author · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2025 FACAM: Design and Optimization of A Compact Energy Efficient FeFET-Based Analog Content Addressable Memory
abstract
Content Addressable Memory (CAM) is known for highly parallel pattern matching capability, which is widely used for data-centric applications and advanced machine learning models that involve associative search tasks. However, most state-of-the-art CAM designs focus on binary/multi-bit CAMs (B/MCAMs) based on CMOS or emerging nonvolatile memories (NVMs), which struggle in scenarios where analog values, rather than discrete levels, need to be stored and searched. Therefore, analog CAMs (ACAMs) offer a promising solution to further increase memory density, improve energy efficiency and extend practical scenarios. Among NVMs, ferroelectric field effect transistors (FeFETs) have emerged as a strong candidate for efficient CAM designs due to the three-terminal structure, high on-off ratio, high OFF resistance and voltage-driven write/read mechanisms. In this paper, we propose FACAM, a compact and energy efficient single-input 2FeFET-1T ACAM cell design, with a two-phase search scheme, that sets the location and width of the matching range through two FeFETs, respectively. We further present a FACAM array which reduces the matchline (ML) voltage swing by shifting ML precharging into the in-cell search operations. We also propose an adaptive scheme to selectively early-terminate second search phase for further search energy optimization. Evaluation results suggest that our proposed FACAM achieves 8.39× and 2.94× energy efficiency compared with the state-of-the-art better 6T-2R ACAM and 2FeFET ACAM. Benchmarking results in deep random forest accelerator show that our approach is 2.14× faster and 7.82× energy efficient than 2FeFET ACAM.
Jiahao Cai, Ann Franchesca Laguna, Thomas Kämpfe, Zheyu Yan, Cheng Zhuo, Xunzhao Yin
ICCAD2
2024 Designing a Lightweight Convolutional Neural Network for Camouflaged Object Detection
abstract
Camouflaged object detection is a challenging task due to the high visual similarity between the object of interest and its surroundings. While deep learning models have shown promising performance, the size and power requirements of most existing models make them unsuitable for deployment in resource-constrained devices. To alleviate this problem, we modified BGNet, a camouflaged object detection network, by replacing its backbone network Res2Net50 with a lighter neural network model such as EfficientNet and MobileNet. Replacing the backbone network with EfficientNetV2-Medium decreased the model size by 1.53× and GPU power consumption by 1.41×. To further reduce the memory footprint, we benchmarked different pruning and quantization algorithms on the resulting network. Our experiments show that applying l2-norm pruning followed by DoReFa quantization reduced the number of multiply-accumulate operations by 3.71×. Our proposed lightweight camouflaged object detection model performs better than the state-of-the-art BGNet, registering weighted F-measure scores of 0.777, 0.739, and 0.808 on CAMO, COD10K, and NC4K, respectively, compared to BGNet with scores of 0.749, 0.722, and 0.788, while also being lighter and requiring lower power.
Mark Edward M. Gonzales, Hans Oswald A. Ibrahim, Elyssia Barrie H. Ong, Ann Franchesca Laguna
COMPSAC4
2024 Design of High-Performance and Compact CAM for Supporting Data-Intensive Applications
abstract
Content addressable memory (CAM) is a special-purpose search engine that can support parallel search directly in memory. CAMs are of increasing interest for machine learning and data analytics applications that require intensive search operations. However, conventional CMOS CAMs have large cell areas and high energy consumption, which limits applicability. Also, many data-intensive applications need more efficient data representation and approximate matching functions, which may not be efficiently realized by conventional ternary CAMs. As such, we introduce a more compact and high-performance CAM design based on non-volatile ferroelectic FET devices. Furthermore, we present a reconfigurable CAM design, MHCAM, to support approximate search for multi-dimensional data. We use DNA alignment as a proxy application to illustrate the design’s application-level benefits.
Liu Liu 0023, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu
ISCAS2
2024 TinyFSL: Tiny Machine Learning for Filipino Sign Language
Loben Klien A Tipan, Alyanna Mari Abalos, Alyana Erin Bondoc, Justin Jarrett To, Joanna Pauline Rivera, Ann Franchesca Laguna, Edward Tighe
PACLIC6
2023 In-Memory Computing Accelerators for Emerging Learning Paradigms
abstract
Over the past decades, emerging, data-driven machine learning (ML) paradigms have increased in popularity, and revolutionized many application domains. To date, a substantial effort has been devoted to devising mechanisms for facilitating the deployment and near ubiquitous use of these memory intensive ML models. This review paper presents the use of in-memory computing (IMC) accelerators for emerging ML paradigms from a bottom-up perspective through the choice of devices, the design of circuits/architectures, to the application-level results.
Dayane Reis, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu
ASP-DAC2
2023 Development of a Computer Vision-Based Road Physical Feature Extraction
abstract
Livability assessment plays a crucial role in urban planning, development, and maintenance. However, the lack of database on the complex physical attributes of city roads of developing countries hinders practical analysis and interpretation of livability indicators. To address this gap, this paper proposes a computer vision-based road physical feature extraction system which aims to automate the collection and annotation of road physical attributes. Our proposed work utilizes modules for lane and object detection to automate data collection. By instating this system, urban planners and policymakers could gain valuable information for making decisions regarding infrastructure and road maintenance. The core object detection of the system achieved an average precision, recall, and mAP of 0.72, 0.46, and 0.52, respectively. For both bike lane detection and road lane counting, the average accuracy achieved is 0.64, and the average mean absolute error achieved is 0.37. Our proposed work provided a computer vision approach to detecting bike lines and counting road lanes capable of storing and visualizing the extracted road physical features.
Juliana Marie Agulto, Yeohan Lorenzo Noroña, Gian Carlo Gutierrez, Eric Benabaye, Geanne Franco, Ann Franchesca Laguna, Jocelynn Cu, Joel P. Ilao, Macario O. Cordel II
IEEE Big Data6
2023 Invited Paper: Algorithm/Hardware Co-Design for Few-Shot Learning at the Edge
abstract
On-device learning is essential to achieve intelligence at the edge, where it is desirable to learn from few samples or even just a single sample. Memory-augmented neural networks (MANNs), which augment neural networks with an attentional memory, can draw on already learnt knowledge patterns and adapt to new but similar tasks. Implementing MANNs on conventional architectures can require a significant amount of costly data transfer, thereby limiting the practical use of MANNs at the edge. In this paper, we introduce algorithm/hardware co-design solutions which exploit compact designs of content addressable memories (CAMs) based on emerging non-volatile memories (e.g., FeFETs) to implement energy-efficient MANN accelerators. The design space of MANN accelerators is systematically analyzed by considering different circuit, architecture, and algorithm options. We further discuss how hyper-dimensional representations of data can be combined with MANNs to overcome the negative effect of device/circuit variabilities on learning quality, thus achieving not only energy-efficient but also accuracy-competitive on-device learning at the edge. We also investigate modeling of device-to-device (D2D) variation in FeFETs using the write-with-verify approach and detail its impact on the energy, delay, and accuracy of the MANN application.
Ann Franchesca Laguna, Mohammad Mehdi Sharifi, Dayane Reis, Liu Liu 0023, Andrew Hennessee, Clayton O'Dell, Ian O'Connor, Michael T. Niemier, Xiaobo Sharon Hu
ICCAD1
2023 A Reconfigurable FeFET Content Addressable Memory for Multi-State Hamming Distance
abstract
Pattern searches, a key operation in many data analytic applications, often deal with data represented by multiple states per dimension. However, hash tables, a common software-based pattern search approach, require a large amount of additional memory, and thus, are limited by the memory wall. A hardware-based solution is to use content-addressable memories (CAMs) that support fast associative searches in parallel. Ternary CAMs (TCAMs) support bit-wise Hamming distance (HD) based searches. Detecting the HD of vectors with multiple states per dimension (i.e., multi-state Hamming distance (MSHD)) can be implemented on TCAMs with one-hot encoding, but requires one TCAM cell per state, leading to a higher area, latency, and energy overhead. We propose a Ferroelectric FET (FeFET)-based multi-state CAM design, MHCAM, which implements MSHD searches in a dense FeFET-based memory array. MHCAM only uses$\lceil log_{2} s \rceil ~2$FeFET CAM cells to represent$s$states or symbols per dimension, and can be reconfigured to 2-bit/4-bit/6-bit/8-bit dimensions. A low-cost sensing circuit with matchline voltage scaling technique is introduced to perform both exact match and threshold match. We use DNA and protein pre-alignment filtering as application case studies to evaluate the application-level benefit of MHCAM. DNA and protein pre-alignment filtering achieve$3.8\times /4.7\times $speedup and$1.7\times /1.8\times $energy improvement compared with the state-of-the-art 2FeFET TCAM-based implementation.
Liu Liu 0023, Ann Franchesca Laguna, Ramin Rajaei, Mohammad Mehdi Sharifi, Arman Kazemi, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu
IEEE Trans. Circuits Syst. I Regul. Pap.2
2022 iMARS: an in-memory-computing architecture for recommendation systems
abstract
Recommendation systems (RecSys) suggest items to users by predicting their preferences based on historical data. Typical RecSys handle large embedding tables and many embedding table related operations. The memory size and bandwidth of the conventional computer architecture restrict the performance of RecSys. This work proposes an in-memory-computing (IMC) architecture (iMARS) for accelerating the filtering and ranking stages of deep neural network-based RecSys. iMARS leverages IMC-friendly embedding tables implemented inside a ferroelectric FET based IMC fabric. Circuit-level and system-level evaluation show that iMARS achieves 16.8x (713x) end-to-end latency (energy) improvement compared to the GPU counterpart for the MovieLens dataset.
Mengyuan Li 0001, Ann Franchesca Laguna, Dayane Reis, Xunzhao Yin, Michael T. Niemier, Xiaobo Sharon Hu
DAC2
2022 Associative Memory Based Experience Replay for Deep Reinforcement Learning
abstract
Experience replay is an essential component in deep reinforcement learning (DRL), which stores the experiences and generates experiences for the agent to learn in real time. Recently, prioritized experience replay (PER) has been proven to be powerful and widely deployed in DRL agents. However, implementing PER on traditional CPU or GPU architectures incurs significant latency overhead due to its frequent and irregular memory accesses. This paper proposes a hardware-software co-design approach to design an associative memory (AM) based PER, AMPER, with an AM-friendly priority sampling operation. AMPER replaces the widely-used time-costly tree-traversal-based priority sampling in PER while preserving the learning performance. Further, we design an in-memory computing hardware architecture based on AM to support AMPER by leveraging parallel in-memory search operations. AMPER shows comparable learning performance while achieving 55× to 270× latency improvement when running on the proposed hardware compared to the state-of-the-art PER running on GPU.
Mengyuan Li 0001, Arman Kazemi, Ann Franchesca Laguna, Xiaobo Sharon Hu
ICCAD3
2022 COSIME: FeFET Based Associative Memory for In-Memory Cosine Similarity Search
abstract
In a number of machine learning models, an input query is searched across the trained class vectors to find the closest feature class vector in cosine similarity metric. However, performing the cosine similarities between the vectors in Von-Neumann machines involves a large number of multiplications, Euclidean normalizations and division operations, thus incurring heavy hardware energy and latency overheads. Moreover, due to the memory wall problem that presents in the conventional architecture, frequent cosine similarity-based searches (CSSs) over the class vectors requires a lot of data movements, limiting the throughput and efficiency of the system. To overcome the aforementioned challenges, this paper introduces COSIME, a general in-memory associative memory (AM) engine based on the ferroelectric FET (FeFET) device for efficient CSS. By leveraging the one-transistor AND gate function of FeFET devices, current-based translinear analog circuit and winner-take-all (WTA) circuitry, COSIME can realize parallel in-memory CSS across all the entries in a memory block, and output the closest word to the input query in cosine similarity metric. Evaluation results at the array level suggest that the proposed COSIME design achieves 333× and 90.5× latency and energy improvements, respectively, and realizes better classification accuracy when compared with an AM design implementing approximated CSS. The proposed in-memory computing fabric is evaluated for an HDC problem, showcasing that COSIME can achieve on average 47.1× and 98.5× speedup and energy efficiency improvements compared with an GPU implementation.
Che-Kai Liu, Haobang Chen, Mohsen Imani, Kai Ni 0004, Arman Kazemi, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu, Liang Zhao 0004, Cheng Zhuo, Xunzhao Yin
ICCAD6
2022 FeFET Multi-Bit Content-Addressable Memories for In-Memory Nearest Neighbor Search
abstract
Nearest neighbor (NN) search computations are at the core of many applications such as few-shot learning, classification, and hyperdimensional computing. As such, efficient hardware support for NN search is highly desired. In-memory computing using emerging devices offers attractive solutions for NN search. Solutions based on ternary content-addressable memories (TCAMs) offer high energy and latency improvements for NN search at the expense of accuracy. In this work, we propose a novel distance function that can be natively evaluated with multi-bit content-addressable memories (MCAMs) based on ferroelectric FETs (FeFETs) to perform a single-step, in-memory NN search. We evaluate the efficacy of FeFET MCAMs in the context of few-shot learning applications with different datasets. As an example, we achieve a 78.54% accuracy for a 5-way, 5-shot classification task for the mini-ImageNet dataset (only 1.5% lower than software-based implementations) when using a 3-bit MCAM for NN search. We consider the effects of FeFET threshold voltage variations on the application accuracy and analyze the area and search energy requirements of FeFET MCAMs for accurate operations. Our results indicate that MCAMs require 2× lower area and search energy than TCAMs to achieve the same accuracy. Furthermore, we experimentally demonstrate a 2-bit implementation of FeFET MCAM using AND arrays from GLOBALFOUNDRIES to further validate the design concept.
Arman Kazemi, Mohammad Mehdi Sharifi, Ann Franchesca Laguna, Franz Müller 0001, Xunzhao Yin, Thomas Kämpfe, Michael T. Niemier, Xiaobo Sharon Hu
IEEE Trans. Computers3
2021 Cross-layer Design for Computing-in-Memory: From Devices, Circuits, to Architectures and Applications
abstract
The era of Big Data, Artificial Intelligence (AI) and Internet of Things (IoT) is approaching, but our underlying computing infrastructures are not sufficiently ready. The end of Moore's law and process scaling as well as the memory wall associated with von Neumann architectures have throttled the rapid development of conventional architectures based on CMOS technology, and cross-layer efforts that involve the interactions from low-end devices to high-end applications have been prominently studied to overcome the aforementioned challenges. On one hand, various emerging devices, e.g., Ferroelectric FET, have been proposed to either sustain the scaling trends or enable novel circuit and architecture innovations. On the other hand, novel computing architectures/algorithms, e.g., computing-in-memory (CiM), have been proposed to address the challenges faced by conventional von Neumann architectures. Naturally, integrated approaches across the emerging devices and computing architectures/algorithms for data-intensive applications are of great interests. This paper uses the FeFET as a representative device, and discuss about the challenges, opportunities and contributions for the emerging trends of cross-layer co-design for CiM.
Hussam Amrouch, Xiaobo Sharon Hu, Mohsen Imani, Ann Franchesca Laguna, Michael T. Niemier, Simon Thomann, Xunzhao Yin, Cheng Zhuo
ASP-DAC4
2021 Attention-in-Memory for Few-Shot Learning with Configurable Ferroelectric FET Arrays
abstract
Attention-in-Memory (AiM), a computing-in-memory (CiM) design, is introduced to implement the attentional layer of Memory Augmented Neural Networks (MANNs). AiM consists of a memory array based on Ferroelectric FETs (FeFET) along with CMOS peripheral circuits implementing configurable functionalities, i.e., it can be dynamically changed from a ternary content-addressable memory (TCAM) to a general-purpose (GP) CiM. When compared to state-of-the art accelerators, AiM achieves comparable end-to-end speed-up and energy for MANNs, with better accuracy (95.14% v.s. 92.21%, and 95.14% v.s. 91.98%) at iso-memory size, for a 5-way 5-shot inference task with the Omniglot dataset.
Dayane Reis, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu
ASP-DAC2
2021 In-Memory Nearest Neighbor Search with FeFET Multi-Bit Content-Addressable Memories
abstract
Nearest neighbor (NN) search is an essential operation in many applications, such as one/few-shot learning and image classification. As such, fast and low-energy hardware support for accurate NN search is highly desirable. Ternary content-addressable memories (TCAMs) have been proposed to accelerate NN search for few-shot learning tasks by implementing$L$∞and Hamming distance metrics, but they cannot achieve software-comparable accuracies. This paper proposes a novel distance function that can be natively evaluated with multi-bit content-addressable memories (MCAMs) based on ferroelectric FETs (Fe-FETs) to perform a single-step, in-memory NN search. Moreover, this approach achieves accuracies comparable to floating-point precision implementations in software for NN classification and one/few-shot learning tasks. As an example, the proposed method achieves a 98.34% accuracy for a 5-way, 5-shot classification task for the Omniglot dataset (only 0.8% lower than software-based implementations) with a 3-bit MCAM. This represents a 13% accuracy improvement over state-of-the-art TCAM-based implementations at iso-energy and iso-delay. The presented distance function is resilient to the effects of FeFET device-to-device variations. Furthermore, this work experimentally demonstrates a 2-bit implementation of FeFET MCAM using AND arrays from GLOBALFOUNDRIES to further validate proof of concept.
Arman Kazemi, Mohammad Mehdi Sharifi, Ann Franchesca Laguna, Franz Müller 0001, Ramin Rajaei, Ricardo Olivo, Thomas Kämpfe, Michael T. Niemier, Xiaobo Sharon Hu
DATE3
2021 In-Memory Computing based Accelerator for Transformer Networks for Long Sequences
abstract
Transformer networks have outperformed recurrent neural networks and convolutional neural networks in various sequential tasks. However, scaling transformer networks for long sequences has been challenging because of memory and compute bottlenecks. Transformer networks are impeded by memory bandwidth limitations because of their low operation per byte ratio resulting in low utilization of GPU's computing resources. In-memory processing can mitigate memory bottlenecks by eliminating the transfer time between memory and compute units. Furthermore, transformer networks use neural attention mechanisms to characterize the relationships between sequence elements. Efficient hardware solutions have been proposed to implement efficient attention mechanisms, which include ternary content addressable memories (TCAM), crossbar arrays (XBars), and processing in-memory (PIM). However, these solutions do not implement a multi-head self-attention mechanism. We propose using a combination of XBars and CAMs to accelerate transformer networks. We improve the speed of transformer networks by (1) computing in-memory, thus minimizing the memory transfer overhead, (2) caching reusable parameters to reduce the number of operations, (3) exploiting the available parallelism in the attention mechanism, and (4) using locality sensitive hashing to filter the number of sequence elements by their importance. Our approach achieves a 200x speedup and 41x energy improvement for a sequence length of 4098.
Ann Franchesca Laguna, Arman Kazemi, Michael T. Niemier, Xiaobo Sharon Hu
DATE1
2021 Exploiting FeFETs via Cross-Layer Design from In-memory Computing Circuits to Meta-Learning Applications
abstract
A ferroelectric FET (FeFET), made by integrating a ferroelectric material layer in the gate stack of a MOSFET, is a device that can behave as both a transistor and a non-volatile storage element. This unique property of FeFETs enables area efficient and low-power merged logic and memory functionality, desirable for many data analytic and machine learning applications. To best exploit this unique feature of FeFETs, cross-layer design practices spanning from circuits and architectures to algorithms and applications is needed. The paper presents FeFET-based circuits and architectures that offer, either independently or in a configurable fashion, content addressable memory (TCAM) and general-purpose compute-in-memory (GP-CiM) functionalities. These in-memory computing modules bring new opportunities to accelerating data-intensive applications. We discuss the use of these FeFET based in-memory computing fabrics in meta-learning applications, specifically as attentional memory. System-level task mapping and end-to-end evaluation will be discussed.
Dayane Reis, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu
DATE2
2021 ICCAD Tutorial Session Paper Ferroelectric FET Technology and Applications: From Devices to Systems
abstract
The rapidly increasing volume and complexity of data is demanding the relentless scaling of computing power. With transistor feature size approaching physical limits, the benefits that CMOS technology can provide is diminishing. For future energy efficient computing systems, researchers aim to exploit various emerging nanotechnologies to replace conventional CMOS technology. In particular, ferroelectric FETs (FeFETs) appear to be a promising candidate to continue improving energy efficiency for data-intensive applications. Advances in FeFET scalability and FeFET compatibility with CMOS have sparked growing interest in device, circuit, and system communities. While FeFET is still evolving, many researchers and developers are already cautiously optimistic about its future. This paper provides a review on FeFET's recent technology advances, challenges, and opportunities, with a particular emphasis upon device modeling and circuit design of FeFET content addressable memory, as well as their applications in machine learning.
Hussam Amrouch, Xiaobo Sharon Hu, Arman Kazemi, Ann Franchesca Laguna, Kai Ni 0004, Michael T. Niemier, Mohammad Mehdi Sharifi, Simon Thomann, Xunzhao Yin, Cheng Zhuo
ICCAD5
2020 Emerging Neural Workloads and Their Impact on Hardware
abstract
We consider existing and emerging neural workloads, and what hardware accelerators might be best suited for said workloads. We begin with a discussion of analog crossbar arrays, which are known to be well-suited for matrix-vector multiplication operations that are commonplace in existing neural network models such as convolutional neural networks (CNNs). We highlight candidate crosspoint devices, what device and materials challenges must be overcome for a given device to be employed in a crossbar array for a computationally interesting neural workload, and how circuit and algorithmic optimizations may be employed to mitigate undesirable characteristics from devices/materials. We then discuss two emerging neural workloads. We first consider machine learning models for one- and few-shot learning tasks (i.e., where a network can be trained with just one or a few, representative examples of a given class). Notably crossbar-based architectures can be used to accelerate said models. Hardware solutions based on content addressable memory arrays will also be discussed. We then consider machine learning models for recommendation systems. Recommendation models, an emerging class of machine learning models, employ distinct neural network architectures that operate of continuous and categorical input features which make hardware acceleration challenging. We will discuss the open research challenges and opportunities within this space.
David Brooks 0001, Martin M. Frank, Tayfun Gokmen, Udit Gupta 0001, Xiaobo Sharon Hu, Shubham Jain 0004, Ann Franchesca Laguna, Michael T. Niemier, Ian O'Connor, Anand Raghunathan, Ashish Ranjan 0001, Dayane Reis, Jacob R. Stevens, Carole-Jean Wu, Xunzhao Yin
DATE7
2020 A Fast and Energy Efficient Computing-in-Memory Architecture for Few-Shot Learning Applications
abstract
Among few-shot learning methods, prototypical networks (PNs) are one of the most popular approaches due to their excellent classification accuracies and network simplicity. Test examples are classified based on their distances from class prototypes. Despite the application-level advantages of PNs, the latency of transferring data from memory to compute units is much higher than the PN computation time. Thus, PNs performance is limited by memory bandwidth. Computing-in-memory addresses this bandwidth-bottleneck problem by bringing a subset of compute units closer to memory. In this work, we propose a CiM-PN framework that enables the computation of distance metrics and prototypes inside the memory. CiM-PN replaces the computationally intensive Euclidean distance metric by the CiM-friendly Manhattan distance metric. Additionally, prototypes are computed using an in-memory mean operation realized by accumulation and division by powers of two, which enables few-shot learning implementations where "shots" are powers of two. The CiM-PN hardware uses CMOS memory cells, as well as CMOS peripherals such as customized sense amplifiers, carry-look-ahead adders, in-place copy buffers and a logarithmic shifter. Compared with a GPU implementation, a CMOS-based CiM-PN achieves speedups of 2808x/111x and energy savings of 2372x/5170x at iso-accuracy for the prototype and nearest-neighbor computation, respectively, and over 2x end-to-end speedup and energy improvements. We also gain 3-14% accuracy improvement when compared to existing non-GPU hardware approaches due to the floating-point CiM operations.
Dayane Reis, Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu
DATE2
2020 Seed-and-Vote based In-Memory Accelerator for DNA Read Mapping
abstract
Genome analysis is becoming more important in the fields of forensic science, medicine, and history. Sequencing technologies such as High Throughput Sequencing (HTS) and Third Generation Sequencing (TGS) have greatly accelerated genome sequencing. However, genome read mapping remains significantly slower than sequencing. Because of the enormous amount of data needed, the speed of the data transfer between the memory and the processing unit limits the execution speed. In-memory computing can help address the memory-bandwidth bottleneck by minimizing data transfers. Ternary Content Addressable Memories (TCAMs) have been used in accelerators because of their fast searching capability for seed-and-extend, a popular read mapping approach. Seed-and-vote, another read mapping approach, is faster than the seed-and-extend approach but has lower accuracies when used with very short reads. Since sequencing technology is moving to longer reads, the seed-and-vote approach is becoming more viable. We propose a genome read mapping accelerator that uses approximate TCAM to execute the Fast Seed and Vote algorithm (FSVA) that can map both short and long reads. We achieved 400X acceleration compared to the seed-and-extend approach BWA-MEM on a CPU and 115X acceleration at 30X energy improvement compared to state-of-the-art in-memory accelerator using the seed-and-extend approach at 98.75% accuracy for 100bp reads.
Ann Franchesca Laguna, Hasindu Gamaarachchi, Xunzhao Yin, Michael T. Niemier, Sri Parameswaran, Xiaobo Sharon Hu
ICCAD1
2019 Design of Hardware-Friendly Memory Enhanced Neural Networks
abstract
Neural networks with external memories have been proven to minimize catastrophic forgetting, a major problem in applications such as lifelong and few-shot learning. However, such memory enhanced neural networks (MENNs) often require a large number of floating point-based cosine distance metric calculations to perform necessary attentional operations, which greatly increases energy consumption and hardware cost. This paper investigates other distance metrics in such neural networks in order to achieve more efficient hardware implementations in MENNs. We propose using content addressable memories (CAMs) to accelerate and simplify attentional operations. Our hardware friendly approach implements fixed point L∞distance calculations via ternary content addressable memories (TCAM) and fixed point L1and L2distance calculations on a general purpose graphical processing unit (GPGPU). As a representative example, a 32-bit floating point-based cosine distance MENN with M · D multiplications has a 99.06% accuracy for the Omniglot 5-way 5-shot classification task. Based on our approach, with just 4-bit fixed point precision, a L∞- L1distance hardware accuracy of 90.35% can be achieved with just 16 TCAM lookups and 16·D addition and subtraction operations. With 4-bit precision and a L∞-L2distance, hardware classification accuracies of 96.00% are possible. Hence, 16 TCAM lookups and 16·D multiplication operations are needed. Assuming the hardware memory has 512 entries, the number of multiplication operations is reduced by 32x versus the cosine distance approach.
Ann Franchesca Laguna, Michael T. Niemier, Xiaobo Sharon Hu
DATE1
2019 Ferroelectric FET Based In-Memory Computing for Few-Shot Learning
abstract
As CMOS technology advances, the performance gap between the CPU and main memory has not improved. Furthermore, the hardware deployed for Internet of Things (IoT) applications need to process ever growing volumes of data, which can further exacerbate the "memory wall". Computing-in-memory (CiM) architectures, where logic and arithmetic operations are performed in memory, can significantly reduce energy and latency overheads associated with data transfer, and potentially alleviate processor-memory bottlenecks. In this paper, we consider the utility of ternary content addressable memory (TCAM) arrays and CiM arrays based on ferroelectric field effect transistors (FeFETs) to support emerging machine learning models that can learn new classes of data with significantly less training overhead - highly desirable in IoT applications. Architecturally, we use TCAM and CiM arrays to implement the external memory module in a memory enhanced neural network (MENN) - which can be used to minimize catastrophic forgetting - a major problem in applications such as lifelong and few-shot learning. As a representative example, we achieve 95.14% accuracy for a few-shot learning task with the Omniglot data set by using a combined L∞ infinity and L1 distance metric computed via a TCAM-CiM cascaded architecture (as opposed to 99.06% accuracy assuming a GPU backed by DRAM). While there is a slight drop in accuracy, the TCAM-CiM approach is 4.34X faster and 4.18X more energy efficient than a CMOS implementation for the same task. The ability of an FeFET to serve as both a compact logic and storage element helps to enable dense CiM and TCAM structures that drive the aforementioned improvements to application-level figures of merit (FOMs).
Ann Franchesca Laguna, Xunzhao Yin, Dayane Reis, Michael T. Niemier, Xiaobo Sharon Hu
ACM Great Lakes Symposium on VLSI1