Abu Sebastian

dblp:15/6452 · DBLP profile ↗
← Back
37ranked-venue papers
1as first author
23since 2021 · last 2025
0000-0001-5603-5243ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 21 · 13 since 2021Artificial intelligence and machine learning · 11 · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 3 since 2021Software engineering, systems software and programming languages · 3 · 2 since 2021Theory of computation · 3 · 2 since 2021Computer networks · 2 · 1 first-author
YearPublicationVenuePosition
2025 On the Expressiveness and Length Generalization of Selective State Space Models on Regular Languages
abstract
Selective state-space models (SSMs) are an emerging alternative to the Transformer, offering the unique advantage of parallel training and sequential inference. Although these models have shown promising performance on a variety of tasks, their formal expressiveness and length generalization properties remain underexplored. In this work, we provide insight into the workings of selective SSMs by analyzing their expressiveness and length generalization performance on regular language tasks, i.e., finite-state automaton (FSA) emulation. We address certain limitations of modern SSM-based architectures by introducing the Selective Dense State-Space Model (SD-SSM), the first selective SSM that exhibits perfect length generalization on a set of various regular language tasks using a single layer. It utilizes a dictionary of dense transition matrices, a softmax selection mechanism that creates a convex combination of dictionary matrices at each time step, and a readout consisting of layer normalization followed by a linear map. We then proceed to evaluate variants of diagonal selective SSMs by considering their empirical performance on commutative and non-commutative automata. We explain the experimental results with theoretical considerations.
Aleksandar Terzic, Michael Hersche, Giacomo Camposampiero, Thomas Hofmann 0001, Abu Sebastian, Abbas Rahimi
AAAI5
2025 Hardware-Aware Compilation and Simulation for In-Memory Computing
abstract
This brief presents an overview of recent tools and research efforts aimed at enhancing the programmability and reliability of In-Memory Computing (IMC)-based systems. We discuss hardware-aware training techniques that improve model resilience to analog device imperfections, and explore mapping strategies that balance accuracy and performance for heterogeneous IMC-based accelerators. Additionally, we examine a compiler framework that abstracts hardware complexities and enables seamless integration of these accelerators into existing deployment pipelines. By combining these approaches with advanced simulation tools, we propose an end-to-end workflow that facilitates the practical deployment and optimization of IMC technologies across diverse memory types and architectural designs.
Asif Ali Khan, Hadjer Benmeziane, Hamid Farzaneh, João Paulo C. de Lima, William Andrew Simon, Yiyu Shi 0001, Zheyu Yan, Abu Sebastian, Xiaobo Sharon Hu, Jerónimo Castrillón, Corey Lammie
CASES8
2025 Live Demonstration: Automated DNN Deployment on the IBM HERMES Project Chip
abstract
For this demonstration, we will showcase the operation of a software stack capable of automatically deploying Matrix-Vector Matrix (MVM) operations of diverse deep learning workloads in a pipelined-manner on a phase-change memory-based analog in-memory computing chip with high-accuracy. For a real chip, each deployment step will be highlighted for a transformer-based network trained to perform an organic chemical reaction prediction task. Additionally, using an emulated mode of operation, these steps will also be highlighted for a Resnet-based network, which has been trained to perform image classification, and a hybrid CNN/LSTM network trained to infer nucleotide sequences from sequences of amplitude values measured from a sequencing device.
Corey Lammie, Julian Büchel, Athanasios Vasilopoulos, Giacomo Camposampiero, Lionel Noussi, William Andrew Simon, Manuel Le Gallo, Abu Sebastian
ISCAS8
2025 NIMA: Near In-Memory High-Precision Accumulation Unit for Heterogeneous Analog/Digital Deep Learning Acceleration
abstract
Analog In-Memory Computing (AIMC) crossbars often face size limitations that hinder mapping entire neural network layers onto a single AIMC tile. To overcome this, tall layers are typically distributed across multiple tiles or within a single tile by multiplexing different sets of weights at various time intervals. However, this generates partial vector-matrix multiplication (VMM) results that need to be accumulated to produce the final output, underscoring the need for integrated accumulation capabilities in AIMC systems. In this work, we propose a Near In-Memory High-Precision Accumulation Unit (NIMA) with built-in internal and external tile accumulation functionality, positioned at the periphery of the AIMC crossbar. This unit leverages the affine correction capabilities of existing systems and ensures high precision during partial VMM accumulation. We have physically implemented NIMA in a 14nm CMOS technology, providing comprehensive performance and area evaluations and comparisons with prior art. We investigate and compare different layer mappings and accumulation schemes supported by the proposed unit, evaluating both latency and Mean Absolute Error (MAE). We further assess the precision of the proposed unit on large layers of ResNet9 and ResNet32 for image classification on CIFAR10/CIFAR100 datasets.
Irem Sanli, Elena Ferro, Athanasios Vasilopoulos, Thomas Boesch, Abu Sebastian, Irem Boybat
ISCAS5
2025 A framework for analog-digital mixed-precision neural network training and inference
abstract
Recent advancements in AI hardware highlight the potential of mixed-signal accelerators, which integrate analog computation for matrix multiplications with reduced-precision digital operations, to achieve superior performance and energy efficiency. In this paper, we present a framework designed to perform hardware-aware training and inference evaluation of neural networks (NNs) on such accelerators. This framework extends an existing toolkit, the IBM Analog Hardware Acceleration Kit (AIHWKit), using a quantization library, enabling flexible layer-wise deployment in either analog or digital units, the latter with configurable precision and quantization options. Our combined framework supports simultaneous quantization-and analog-aware training as well as post-training calibration routines. It can also evaluate the accuracy of NNs when deployed on mixed-signal accelerators. We demonstrate the need of such a framework through ablation studies on a ResNet-based vision model and a BERT-based language model, highlighting the importance of its functionality for maximizing accuracy during deployment. Our contribution is open-sourced as part of the core code of AIHWKit [1].
Athanasios Vasilopoulos, Emma Boulharts, Corey Lammie, Julian Büchel, Hadjer Benmeziane, Manuel Le Gallo, Abu Sebastian
ISCAS7
2025 Can Large Reasoning Models do Analogical Reasoning under Perceptual Uncertainty?
abstract
This work presents a first evaluation of two state-of-the-art Large Reasoning Models (LRMs), OpenAI’s o3-mini and DeepSeek R1, on analogical reasoning, focusing on well-established nonverbal human IQ tests based on Raven’s progressive matrices. We benchmark with the I-RAVEN dataset and its extension, I-RAVEN-X, which tests the ability to generalize to longer reasoning rules and ranges of the attribute values. To assess the influence of visual uncertainties on these symbolic analogical reasoning tests, we extend the I-RAVEN-X dataset, which otherwise assumes an oracle perception. We adopt a two-fold strategy to simulate this imperfect visual perception: 1) we introduce confounding attributes which, being sampled at random, do not contribute to the prediction of the correct answer of the puzzles, and 2) smooth the distributions of the input attributes’ values. We observe a sharp decline in OpenAI’s o3-mini task accuracy, dropping from 86.6% on the original I-RAVEN to just 17.0%—approaching random chance—on the more challenging I-RAVEN-X, which increases input length and range and emulates perceptual uncertainty. This drop occurred despite spending 3.4x more reasoning tokens. A similar trend is also observed for DeepSeek R1: from 80.6% to 23.2%. On the other hand, a neuro-symbolic probabilistic abductive model, ARLC, that achieves state-of-the-art performances on I-RAVEN, can robustly reason under all these out-of-distribution tests, maintaining strong accuracy with only a modest accuracy reduction from 98.6% to 88.0%. Our code is available at https://github.com/IBM/raven-large-language-models.
Giacomo Camposampiero, Michael Hersche, Roger Wattenhofer, Abu Sebastian, Abbas Rahimi
NeSy4
2025 Analog Foundation Models
abstract
Analog in-memory computing (AIMC) is a promising compute paradigm to improve speed and power efficiency of neural network inference beyond the limits of conventional von Neumann-based architectures. However, AIMC introduces fundamental challenges such as noisy computations and strict constraints on input and output quantization. Because of these constraints and imprecisions, off-the-shelf LLMs are not able to achieve 4-bit-level performance when deployed on AIMC-based hardware. While researchers previously investigated recovering this accuracy gap on small, mostly vision-based models, a generic method applicable to LLMs pre-trained on trillions of tokens does not yet exist. In this work, we introduce a general and scalable method to robustly adapt LLMs for execution on noisy, low-precision analog hardware. Our approach enables state-of-the-art models — including Phi-3-mini-4k-instruct and Llama-3.2-1B-Instruct — to retain performance comparable to 4-bit weight, 8-bit activation baselines, despite the presence of analog noise and quantization constraints. Additionally, we show that as a byproduct of our training methodology, analog foundation models can be quantized for inference on low-precision digital hardware. Finally, we show that our models also benefit from test-time compute scaling, showing better scaling behavior than models trained with 4-bit weight and 8-bit static input quantization. Our work bridges the gap between high-capacity LLMs and efficient analog hardware, offering a path toward energy-efficient foundation models. Code is available at [github.com/IBM/analog-foundation-models](https://github.com/IBM/analog-foundation-models).
Julian Büchel, Iason Chalas, Giovanni Acampa, An Chen 0002, Omobayode Fagbohungbe, Hsinyu Tsai, Kaoutar El Maghraoui, Manuel Le Gallo, Abbas Rahimi, Abu Sebastian
NeurIPS10
2025 CiMBA: Accelerating Genome Sequencing Through On-Device Basecalling via Compute-in-Memory
abstract
As genome sequencing is finding utility in a wide variety of domains beyond the confines of traditional medical settings, its computational pipeline faces two significant challenges. First, the creation of up to 0.5 GB of data per minute imposes substantial communication and storage overheads. Second, the sequencing pipeline is bottlenecked at the basecalling step, consuming >40% of genome analysis time. A range of proposals have attempted to address these challenges, with limited success. We propose to address these challenges with a Compute-in-Memory Basecalling Accelerator (CiMBA), the first embedded ($\sim 25$mm$^{2}$) accelerator capable of real-time, on-device basecalling, coupled with AnaLog (AL)-Dorado, a new family of analog focused basecalling DNNs. Our resulting hardware/software co-design greatly reduces data communication overhead, is capable of a throughput of 4.77 million bases per second, 24× that required for real-time operation, and achieves 17 × /27× power/area efficiency over the best prior basecalling embedded accelerator while maintaining a high accuracy comparable to state-of-the-art software basecallers.
William Andrew Simon, Irem Boybat, Riselda Kodra, Elena Ferro, Gagandeep Singh 0002, Mohammed Alser, Shubham Jain 0004, Hsinyu Tsai, Geoffrey W. Burr, Onur Mutlu, Abu Sebastian
IEEE Trans. Parallel Distributed Syst.11
2024 Analog AI as a Service: A Cloud Platform for In-Memory Computing
abstract
This paper introduces the Analog AI Cloud Composer platform, a service that allows users to access Analog In-Memory Computing (AIMC) simulation and computing resources over the cloud. We introduce the concept of an Analog AI as a Service (AAaaS). AIMC offers a novel approach for decreasing both the latency and energy usage associated with Deep Neural Network (DNN) inference and training. This platform democratizes access to AIMC computing, making it available to a broader audience, including researchers, developers, and businesses. Emphasizing a user-friendly, no-code approach, AAaaS integrates the Analog Hardware Acceleration Kit (AIHWKit) simulation platform within a fully managed cloud environment. We discuss the architecture of the Analog AI Cloud Composer (AAICC), focusing on its key services such as inference, training, and AIMC hardware access. The platform's design, grounded in cloud services and guidelines, ensures a secure, data-centric user experience with robust control and validation mechanisms.
Kaoutar El Maghraoui, Kim Tran, Kurtis Ruby, Borja Godoy, Jordan Murray, Manuel Le Gallo-Bourdeau, Todd Deshane, Pablo Gonzalez, Diego Moreda, Hadjer Benmeziane, Corey Lammie, Julian Büchel, Malte J. Rasch, Abu Sebastian, Vijay Narayanan
SSE14
2024 Zero-Shot Classification Using Hyperdimensional Computing
abstract
Classification based on Zero-shot Learning (ZSL) is the ability of a model to classify inputs into novel classes on which the model has not previously seen any training examples. Providing a set of attributes associated with the new class as an auxiliary descriptor is one of the favored approaches to solving this challenging task. In this work, inspired by Hyperdimensional Computing (HDC), we propose the use of stationary distributed binary codebooks in an attribute encoder to compactly represent a computationally simple end-to-end trainable model, which we name Hyperdimensional Computing Zero-shot Classifier (HDC-ZSC). It additionally consists of a trainable image encoder, and a similarity kernel. HDC-ZSC achieves Pareto optimal results with a 63.8 % top-1 classification accuracy on the CUB-200 dataset by having only 26.6 million trainable parameters. Compared to two other state-of-the-art non-generative approaches, HDC-ZSC achieves 4.3% and 9.9% better accuracy, while they require more than 1.85× and 1.72× parameters compared to HDC-ZSC, respectively.
Samuele Ruffino, Geethan Karunaratne, Michael Hersche, Luca Benini, Abu Sebastian, Abbas Rahimi
DATE5
2024 RETRO-LI: Small-Scale Retrieval Augmented Generation Supporting Noisy Similarity Searches and Domain Shift Generalization
abstract
The retrieval augmented generation (RAG) system such as RETRO has been shown to improve language modeling capabilities and reduce toxicity and hallucinations by retrieving from a database of non-parametric memory containing trillions of entries. We introduce RETRO-LI that shows retrieval can also help using a small scale database, but it demands more accurate and better neighbors when searching in a smaller hence sparser non-parametric memory. This can be met by using a proper semantic similarity search. We further propose adding a regularization to the non-parametric memory for the first time: it significantly reduces perplexity when the neighbor search operations are noisy during inference, and it improves generalization when a domain shift occurs. We also show that the RETRO-LI’s non-parametric memory can potentially be implemented on analog in-memory computing hardware, exhibiting O(1) search time while causing noise in retrieving neighbors, with minimal (<1%) performance loss. Our code is available at: https://github.com/IBM/Retrieval-Enhanced-Transformer-Little
Gentiana Rashiti, Geethan Karunaratne, Mrinmaya Sachan, Abu Sebastian, Abbas Rahimi
ECAI4
2024 A Precision-Optimized Fixed-Point Near-Memory Digital Processing Unit for Analog In-Memory Computing
abstract
Analog In-Memory Computing (AIMC) is an emerging technology for fast and energy-efficient Deep Learning (DL) inference. However, a certain amount of digital post-processing is required to deal with circuit mismatches and non-idealities associated with the memory devices. Efficient near-memory digital logic is critical to retain the high area/energy efficiency and low latency of AIMC. Existing systems adopt Floating Point 16 (FP16) arithmetic with limited parallelization capability and high latency. To overcome these limitations, we propose a Near-Memory digital Processing Unit (NMPU) based on fixed-point arithmetic. It achieves competitive accuracy and higher computing throughput than previous approaches while minimizing the area overhead. Moreover, the NMPU supports standard DL activation steps, such as ReLU and Batch Normalization. We perform a physical implementation of the NMPU design in a 14 nm CMOS technology and provide detailed performance, power, and area assessments. We validate the efficacy of the NMPU by using data from an AIMC chip and demonstrate that a simulated AIMC system with the proposed NMPU outperforms existing FP16- based implementations, providing 139 × speed-up, 7.8 × smaller area, and a competitive power consumption. Additionally, our approach achieves an inference accuracy of 86.65 %/65.06 %, with an accuracy drop of just 0.12 %/0.4 % compared to the FP16 baseline when benchmarked with ResNet9/ResNet32 networks trained on the CIFAR10/CIFAR100 datasets, respectively.
Elena Ferro, Athanasios Vasilopoulos, Corey Lammie, Manuel Le Gallo, Luca Benini, Irem Boybat, Abu Sebastian
ISCAS7
2024 Improving the Accuracy of Analog-Based In-Memory Computing Accelerators Post-Training
abstract
Analog-Based In-Memory Computing (AIMC) inference accelerators can be used to efficiently execute Deep Neural Network (DNN) inference workloads. However, to mitigate accuracy losses, due to circuit and device non-idealities, Hardware-Aware (HWA) training methodologies must be employed. These typically require significant information about the underlying hardware. In this paper, we propose two Post-Training (PT) optimization methods to improve accuracy after training is performed. For each crossbar, the first optimizes the conductance range of each column, and the second optimizes the input, i.e, Digital-to-Analog Converter (DAC), range. It is demonstrated that, when these methods are employed, the complexity during training, and the amount of information about the underlying hardware can be reduced, with no notable change in accuracy (≤0.1%) when finetuning the pretrained RoBERTa transformer model for all General Language Understanding Evaluation (GLUE) benchmark tasks. Additionally, it is demonstrated that further optimizing learned parameters PT improves accuracy.
Corey Lammie, Athanasios Vasilopoulos, Julian Büchel, Giacomo Camposampiero, Manuel Le Gallo, Malte J. Rasch, Abu Sebastian
ISCAS7
2024 Towards Learning Abductive Reasoning Using VSA Distributed Representations
Giacomo Camposampiero, Michael Hersche, Aleksandar Terzic, Roger Wattenhofer, Abu Sebastian, Abbas Rahimi
NeSy (1)5
2023 MIMONets: Multiple-Input-Multiple-Output Neural Networks Exploiting Computation in Superposition
abstract
With the advent of deep learning, progressively larger neural networks have been designed to solve complex tasks. We take advantage of these capacity-rich models to lower the cost of inference by exploiting computation in superposition. To reduce the computational burden per input, we propose Multiple-Input-Multiple-Output Neural Networks (MIMONets) capable of handling many inputs at once. MIMONets augment various deep neural network architectures with variable binding mechanisms to represent an arbitrary number of inputs in a compositional data structure via fixed-width distributed representations. Accordingly, MIMONets adapt nonlinear neural transformations to process the data structure holistically, leading to a speedup nearly proportional to the number of superposed input items in the data structure. After processing in superposition, an unbinding mechanism recovers each transformed input of interest. MIMONets also provide a dynamic trade-off between accuracy and throughput by an instantaneous on-demand switching between a set of accuracy-throughput operating points, yet within a single set of fixed parameters. We apply the concept of MIMONets to both CNN and Transformer architectures resulting in MIMOConv and MIMOFormer, respectively. Empirical evaluations show that MIMOConv achieves $\approx 2$–$4\times$ speedup at an accuracy delta within [+0.68, -3.18]% compared to WideResNet CNNs on CIFAR10 and CIFAR100. Similarly, MIMOFormer can handle $2$–$4$ inputs at once while maintaining a high average accuracy within a [-1.07, -3.43]% delta on the long range arena benchmark. Finally, we provide mathematical bounds on the interference between superposition channels in MIMOFormer. Our code is available at https://github.com/IBM/multiple-input-multiple-output-nets.
Nicolas Menet, Michael Hersche, Geethan Karunaratne, Luca Benini, Abu Sebastian, Abbas Rahimi
NeurIPS5
2023 ALPINE: Analog In-Memory Acceleration With Tight Processor Integration for Deep Learning
abstract
Analog in-memory computing (AIMC) cores offers significant performance and energy benefits for neural network inference with respect to digital logic (e.g., CPUs). AIMCs accelerate matrix-vector multiplications, which dominate these applications' run-time. However, AIMC-centric platforms lack the flexibility of general-purpose systems, as they often have hard-coded data flows and can only support a limited set of processing functions. With the goal of bridging this gap in flexibility, we present a novel system architecture that tightly integrates analog in-memory computing accelerators into multi-core CPUs in general-purpose systems. We developed a powerful gem5-based full system-level simulation framework into the gem5-X simulator, ALPINE, which enables an in-depth characterization of the proposed architecture. ALPINE allows the simulation of the entire computer architecture stack from major hardware components to their interactions with the Linux OS. Within ALPINE, we have defined a custom ISA extension and a software library to facilitate the deployment of inference models. We showcase and analyze a variety of mappings of different neural network types, and demonstrate up to 20.5x/20.8x performance/energy gains with respect to a SIMD-enabled ARM CPU implementation for convolutional neural networks, multi-layer perceptrons, and recurrent neural networks.
Joshua Alexander Harrison Klein, Irem Boybat, Yasir Mahmood Qureshi, Martino Dazzi, Alexandre Levisse, Giovanni Ansaloni, Marina Zapater, Abu Sebastian, David Atienza 0001
IEEE Trans. Computers8
2023 Generalized Key-Value Memory to Flexibly Adjust Redundancy in Memory-Augmented Networks
abstract
Memory-augmented neural networks enhance a neural network with an external key-value (KV) memory whose complexity is typically dominated by the number of support vectors in the key memory. We propose a generalized KV memory that decouples its dimension from the number of support vectors by introducing a free parameter that can arbitrarily add or remove redundancy to the key memory representation. In effect, it provides an additional degree of freedom to flexibly control the tradeoff between robustness and the resources required to store and compute the generalized KV memory. This is particularly useful for realizing the key memory on in-memory computing hardware where it exploits nonideal, but extremely efficient nonvolatile memory devices for dense storage and computation. Experimental results show that adapting this parameter on demand effectively mitigates up to 44% nonidealities, at equal accuracy and number of devices, without any need for neural network retraining.
Denis Kleyko, Geethan Karunaratne, Jan M. Rabaey, Abu Sebastian, Abbas Rahimi
IEEE Trans. Neural Networks Learn. Syst.4
2022 Constrained Few-shot Class-incremental Learning
abstract
Continually learning new classes from fresh data without forgetting previous knowledge of old classes is a very challenging research problem. Moreover, it is imperative that such learning must respect certain memory and computational constraints such as (i) training samples are limited to only a few per class, (ii) the computational cost of learning a novel class remains constant, and (iii) the memory footprint of the model grows at most linearly with the number of classes observed. To meet the above constraints, we propose C-FSCIL, which is architecturally composed of a frozen meta-learned feature extractor, a trainable fixed-size fully connected layer, and a rewritable dynamically growing memory that stores as many vectors as the number of encountered classes. C-FSCIL provides three update modes that offer a trade-off between accuracy and compute-memory cost of learning novel classes. C-FSCIL exploits hyperdimensional embedding that allows to continually express many more classes than the fixed dimensions in the vector space, with minimal interference. The quality of class vector representations is further improved by aligning them quasi-orthogonally to each other by means of novel loss functions. Experiments on the CIFAR100, mini-ImageNet, and Omniglot datasets show that C-FSCIL outperforms the baselines with remarkable accuracy and compression. It also scales up to the largest problem size ever tried in this few-shot setting by learning 423 novel classes on top of 1200 base classes with less than 1.6% accuracy drop. Our code is available at https://github.com/IBM/constrained-FSCIL.
Michael Hersche, Geethan Karunaratne, Giovanni Cherubini, Luca Benini, Abu Sebastian, Abbas Rahimi
CVPR5
2022 Wireless On-Chip Communications for Scalable In-memory Hyperdimensional Computing
abstract
Hyperdimensional computing (HDC) is an emerging computing paradigm that represents, manipulates, and communicates data using very long random vectors (aka hypervectors). Among different hardware platforms capable of executing HDC algorithms, in-memory computing (IMC) systems have been recently proved to be one of the most energy-efficient options, due to hypervector manipulations in the memory itself that reduces data movement. Although implementations of HDC on single IMC cores have been made, their parallelization is still unresolved due to the communication challenges that these novel architectures impose and that traditional Networks-on-Chip and Networks-in-Package were not designed for. To cope with this difficulty, we propose the use of wireless on-chip communication technology in unique ways. We are particularly interested in physically distributing a large number of IMC cores performing similarity search across a chip, and maintaining the classification accuracy when each of which is queried with a slightly different version of a bundled hypervector. To achieve it, we introduce a novel over-the-air computing that consists of defining different binary decision regions in the receivers so as to compute the logical majority operation (i.e., bundling, or superposition) required in HDC. It introduces moderate overheads of a single antenna and receiver per IMC core. By doing so, we achieve a joint broadcast distribution and computation with a performance and efficiency unattainable with wired interconnects, which in turn enables massive parallelization of the architecture. It is demonstrated that the proposed approach allows to both bundle at least three hypervectors and scale similarity search to 64 IMC cores seamlessly, while incurring an average bit error ratio of 0.01 without any impact in the accuracy of a generic HDC-based classifier working with 512-bit vectors.
Robert Guirado, Abbas Rahimi, Geethan Karunaratne, Eduard Alarcón, Abu Sebastian, Sergi Abadal
IJCNN5
2022 MNEMOSENE: Tile Architecture and Simulator for Memristor-based Computation-in-memory
abstract
In recent years, we are witnessing a trend toward in-memory computing for future generations of computers that differs from traditional von-Neumann architecture in which there is a clear distinction between computing and memory units. Considering that data movements between the central processing unit (CPU) and memory consume several orders of magnitude more energy compared to simple arithmetic operations in the CPU, in-memory computing will lead to huge energy savings as data no longer needs to be moved around between these units. In an initial step toward this goal, new non-volatile memory technologies, e.g., resistive RAM (ReRAM) and phase-change memory (PCM), are being explored. This has led to a large body of research that mainly focuses on the design of the memory array and its peripheral circuitry. In this article, we mainly focus on the tile architecture (comprising a memory array and peripheral circuitry) in which storage and compute operations are performed in the (analog) memory array and the results are produced in the (digital) periphery. Such an architecture is termed compute-in-memory-periphery (CIM-P). More precisely, we derive an abstract CIM-tile architecture and define its main building blocks. To bridge the gap between higher-level programming languages and the underlying (analog) circuit designs, an instruction-set architecture is defined that is intended to control and, in turn, sequence the operations within this CIM tile to perform higher-level more complex operations. Moreover, we define a procedure to pipeline the CIM-tile operations to further improve the performance. To simulate the tile and perform design space exploration considering different technologies and parameters, we introduce the fully parameterized first-of-its-kind CIM tile simulator and compiler. Furthermore, the compiler is technology-aware when scheduling the CIM-tile instructions. Finally, using the simulator, we perform several preliminary design space explorations regarding the three competing technologies, ReRAM, PCM, and STT-MRAM concerning CIM-tile parameters, e.g., the number of ADCs. Additionally, we investigate the effect of pipelining in relation to the clock speeds of the digital periphery assuming the three technologies. In the end, we demonstrate that our simulator is also capable of reporting energy consumption for each building block within the CIM tile after the execution of in-memory kernels considering the data-dependency on the energy consumption of the memory array. All the source codes are publicly available.
Mahdi Zahedi, Muath Abu Lebdeh, Christopher Bengel, Dirk J. Wouters, Stephan Menzel, Manuel Le Gallo, Abu Sebastian, Stephan Wong, Said Hamdioui
ACM J. Emerg. Technol. Comput. Syst.7
2021 Architecting more than Moore: wireless plasticity for massive heterogeneous computer architectures (WiPLASH)
abstract
This paper presents the research directions pursued by the WiPLASH European project, pioneering on-chip wireless communications as a disruptive enabler towards next-generation computing systems for artificial intelligence (AI). We illustrate the holistic approach driving our research efforts, which encompass expertises and abstraction levels ranging from physical design of embedded graphene antennas to system-level evaluation of wirelessly-communicating heterogeneous systems.
Joshua Alexander Harrison Klein, Alexandre Levisse, Giovanni Ansaloni, David Atienza 0001, Marina Zapater, Martino Dazzi, Geethan Karunaratne, Irem Boybat, Abu Sebastian, Davide Rossi 0001, Francesco Conti 0001, Elana Pereira de Santana, Peter Haring Bolívar, Mohamed Saeed, Renato Negra, Kun-Ta Wang, Max Christian Lemme, Akshay Jain 0001, Robert Guirado, Hamidreza Taghvaee, Sergi Abadal
CF9
2021 Accurate Weight Mapping in a Multi-Memristive Synaptic Unit
abstract
In-memory computing using memristive devices is a promising non-von Neumann approach for making energy- efficient deep learning inference hardware. Synaptic units comprising one or more memristive devices organized in a crossbar configuration are capable of performing the matrix-vector multiply operations in place by exploiting the Kirchhoff's circuits laws. In this paper, we propose a weight mapping algorithm to efficiently program such a synaptic unit comprising multiple phase change memory (PCM) devices to target conductance values. To evaluate the programming scheme, a simulator based on the measured programming characteristics of 10,000 PCM devices is developed. It is shown that the synaptic unit can be programmed reliably without significant overhead in programming time or energy compared to a unit comprising a single PCM device, while gaining resilience to device-level non-idealities and yield. The algorithm is experimentally verified on a prototype PCM unit cell fabricated in the 90nm CMOS technology node.
Michele Martemucci, Benedikt Kersting, Riduan Khaddam-Aljameh, Irem Boybat, S. R. Nandakumar, Urs Egger, M. J. BrightSky, Robert L. Bruce, Manuel Le Gallo, Abu Sebastian
ISCAS10
2021 Efficient Pipelined Execution of CNNs Based on In-Memory Computing and Graph Homomorphism Verification
abstract
In-memory computing is an emerging computing paradigm enabling deep-learning inference at significantly higher energy-efficiency and reduced latency. The essential idea is mapping the synaptic weights of each layer to one or more in-memory computing (IMC) cores. During inference, these cores perform the associated matrix-vector multiplications in place with O(1) time complexity, obviating the need to move the synaptic weights to additional processing units. Moreover, this architecture enables the execution of these networks in a highly pipelined fashion. However, a key challenge is designing an efficient communication fabric for the IMC cores. In this work, we present one such communication fabric based on a graph topology that is well-suited for the widely successful convolutional neural networks (CNNs). We show that this communication fabric facilitates the pipelined execution of all state-of-the-art CNNs by proving the existence of a homomorphism between the graph representations of these networks and that corresponding to the proposed communication fabric. We then present a quantitative comparison with established communication topologies and show that our proposed topology achieves the lowest bandwidth requirements per communication channel. Finally, we present one hardware implementation and show a concrete example of mapping ResNet-32 onto an IMC core array interconnected via the proposed communication fabric.
Martino Dazzi, Abu Sebastian, Thomas P. Parnell, Pier Andrea Francese, Luca Benini, Evangelos Eleftheriou
IEEE Trans. Computers2
2020 ESSOP: Efficient and Scalable Stochastic Outer Product Architecture for Deep Learning
abstract
Deep neural networks (DNNs) have surpassed human-level accuracy in a variety of cognitive tasks but at the cost of significant memory/time requirements in DNN training. This limits their deployment in energy and memory limited applications that require real-time learning. Matrix-vector multiplications (MVM) and vector-vector outer product (VVOP) are the two most expensive operations associated with training of DNNs. Strategies to improve the efficiency of MVM computation in hardware have been demonstrated with minimal impact on training accuracy. However, the VVOP computation remains a relatively less explored bottleneck even with the aforementioned strategies. Stochastic computing (SC) has been proposed to improve the efficiency of VVOP computation but on relatively shallow networks with bounded activation functions and floatingpoint (FP) scaling of activation gradients. In this paper, we propose ESSOP, an efficient and scalable stochastic outer product architecture based on the SC paradigm. We introduce efficient techniques to generalize SC for weight update computation in DNNs with the unbounded activation functions (e.g., ReLU), required by many state-of-the-art networks. Our architecture reduces the computational cost by re-using random numbers and replacing certain FP multiplication operations by bit shift scaling. We show that the ResNet-32 network with 33 convolution layers and a fully-connected layer can be trained with ESSOP on the CIFAR-10 dataset to achieve baseline comparable accuracy. Hardware design of ESSOP at 14nm technology node shows that, compared to a highly pipelined FP16 multiplier design, ESSOP is 82.2% and 93.7% better in energy and area efficiency respectively for outer product computation.
Vinay Joshi, Geethan Karunaratne, Manuel Le Gallo, Irem Boybat, Christophe Piveteau, Abu Sebastian, Bipin Rajendran, Evangelos Eleftheriou
ISCAS6
2020 Accurate Emulation of Memristive Crossbar Arrays for In-Memory Computing
abstract
In-memory computing is an emerging non-von Neumann computing paradigm where certain computational tasks are performed in memory by exploiting the physical attributes of the memory devices. Memristive devices such as phase-change memory (PCM), where information is stored in terms of their conductance levels, are especially well suited for in-memory computing. In particular, memristive devices, when organized in a crossbar configuration can be used to perform matrix-vector multiply operations by exploiting Kirchhoff's circuit laws. To explore the feasibility of such in-memory computing cores in applications such as deep learning as well as for system-level architectural exploration, it is highly desirable to develop an accurate hardware emulator that captures the key physical attributes of the memristive devices. Here, we present one such emulator for PCM and experimentally validate it using measurements from a PCM prototype chip. Moreover, we present an application of the emulator for neural network inference where our emulator can capture the conductance evolution of approximately 400,000 PCM devices remarkably well.
Anastasios Petropoulos, Irem Boybat, Manuel Le Gallo, Evangelos Eleftheriou, Abu Sebastian, Theodore Antonakopoulos 0001
ISCAS5
2020 File Classification Based on Spiking Neural Networks
abstract
In this paper, we propose a system for file classification in large data sets based on spiking neural networks (SNNs). File information contained in key-value metadata pairs is mapped by a novel correlative temporal encoding scheme to spike patterns that are input to an SNN. The correlation between input spike patterns is determined by a file similarity measure. Unsupervised training of such networks using spike-timing-dependent plasticity (STDP) is addressed first. Then, supervised SNN training is considered by backpropagation of an error signal that is obtained by comparing the spike pattern at the output neurons with a target pattern representing the desired class. The classification accuracy is measured for various publicly available data sets with tens of thousands of elements, and compared with other learning algorithms, including logistic regression and support-vector machines. Simulation results indicate that the proposed SNN-based system using memristive synapses may represent a valid alternative to classical machine learning algorithms for inference tasks, especially in environments with asynchronous ingest of input data and limited resources.
Ana Stanojevic, Giovanni Cherubini, Timoleon Moraitis, Abu Sebastian
ISCAS4
2019 Applications of Computation-In-Memory Architectures based on Memristive Devices
abstract
Today's computing architectures and device technologies are unable to meet the increasingly stringent demands on energy and performance posed by emerging applications. Therefore, alternative computing architectures are being explored that leverage novel post-CMOS device technologies. One of these is a Computation-in-Memory architecture based on memristive devices. This paper describes the concept of such an architecture and shows different applications that could significantly benefit from it. For each application, the algorithm, the architecture, the primitive operations, and the potential benefits are presented. The applications cover the domains of data analytics, signal processing, and machine learning.
Said Hamdioui, Hoang Anh Du Nguyen, Mottaqiallah Taouil, Abu Sebastian, Manuel Le Gallo, Sandeep Pande, Siebren Schaafsma, Francky Catthoor, Shidhartha Das, Fernando García-Redondo, Geethan Karunaratne, Abbas Rahimi, Luca Benini
DATE4
2019 Multi-ReRAM Synapses for Artificial Neural Network Training
abstract
Metal-oxide-based resistive memory devices (ReRAM) are being actively researched as synaptic elements of neuromorphic co-processors for training deep neural networks (DNNs). However, device-level non-idealities are posing significant challenges. In this work we present a multi-ReRAM-based synaptic architecture with a counter-based arbitration scheme that shows significant promise. We present a 32×2 crossbar array comprising Pt/HfO2/Ti/TiN-based ReRAM devices with multi-level storage capability and bidirectional conductance response. We study the device characteristics in detail and model the conductance response. We show through simulations that an in-situ trained DNN with a multi-ReRAM synaptic architecture can perform handwritten digit classification task with high accuracies, only 2% lower than software simulations using floating point precision, despite the stochasticity, nonlinearity and large conductance change granularity associated with the devices. Moreover, we show that a network can achieve accuracies > 80% even with just binary ReRAM devices with this architecture.
Irem Boybat, Cecilia Giovinazzo, Elmira Shahrabi, Igor Krawczuk, Iason Giannopoulos, Christophe Piveteau, Manuel Le Gallo, Carlo Ricciardi, Abu Sebastian, Evangelos Eleftheriou, Yusuf Leblebici
ISCAS9
2018 Spiking Neural Networks Enable Two-Dimensional Neurons and Unsupervised Multi-Timescale Learning
abstract
The capabilities of artificial neural networks (ANNs) are limited by the operations possible at their individual neurons and synapses. For instance, each neuron's activation only represents a single scalar variable. In addition, because neuronal activations may be dominated by a single timescale in the synaptic input, unsupervised learning from data with multiple timescales has not been generally possible. Here we address these by exploiting the continuous-time and asynchronous operation of spiking neural networks (SNNs), i.e. a biologically-inspired type of ANNs. First, we demonstrate how input neurons can be two-dimensional (2D), i.e. each represent two variables. Second, we show unsupervised learning from multiple timescales simultaneously. 2D neurons operate by allocating each variable to a different timescale in their activation, i.e. one variable corresponds to the timing of individual spikes, and another to the spike rate. We show how these can be modulated separately but simultaneously, and we apply this mixed coding technique to encoding images with two modalities, namely, colour and brightness. Unsupervised multi-timescale learning is achieved by synapses with spike-timing-dependent plasticity, combined with varying degrees of short-term plasticity. We demonstrate the successful application of this learning scheme on the unsupervised classification of bimodal pictures encoded by our 2D neurons. Taken together, our results show that SNNs are capable of increasing both the information content of each neuron and the exploitable data in the input. We suggest that through these unique features, SNNs may increase the performance and broaden the applicability of ANNs.
Timoleon Moraitis, Abu Sebastian, Evangelos Eleftheriou
IJCNN2
2018 Mixed-precision architecture based on computational memory for training deep neural networks
abstract
Deep neural networks (DNN) have revolutionized the field of machine learning by providing unprecedented human-like performance in solving many real-world problems such as image or speech recognition. Training of large DNNs, however, is a computationally intensive task, and this necessitates the development of novel computing architectures targeting this application. A computational memory unit where resistive memory devices are organized in crossbar arrays can be used to store the synaptic weights in their conductance states. The expensive multiply accumulate operations can be performed in place using Kirchhoff's circuit laws in a non-von Neumann manner. However, a key challenge remains the inability to alter the conductance states of the devices in a reliable manner during the weight update process. We propose a mixed-precision architecture that combines a computational memory unit storing the synaptic weights with a digital processing unit and an additional memory unit that stores the accumulated weight updates in high precision. The new architecture delivers classification accuracies comparable to those of floating-point implementations without being constrained by challenges associated with the non-ideal weight update characteristics of emerging resistive memories. The computational memory unit in a two layer neural network realized using nonlinear stochastic models of phase-change memory achieves a test accuracy of 97.40% in the MNIST digit classification problem.
S. R. Nandakumar, Manuel Le Gallo, Irem Boybat, Bipin Rajendran, Abu Sebastian, Evangelos Eleftheriou
ISCAS5
2018 Exploiting the non-linear current-voltage characteristics for resistive memory readout
abstract
Various resistive memory technologies are finding application in the space of storage-class memory and emerging non-von Neumann computing systems. For both applications, a key enabling technology is the ability to store multiple resistance levels in a single memory cell. The resistance states of these devices are typically measured in the low-field regime, where the electrical transport can be assumed to be Ohmic. However, when biased at slightly higher voltages, they exhibit significantly nonlinear I-V characteristics. In this paper, we demonstrate how this field dependence of the resistance values can be exploited in various applications. We present simulation and experimental results where readout schemes based on the non-linear I-V behavior are used to enhance the readout margin and also to compensate for resistance drift.
Nikolaos Papandreou, Abu Sebastian, Haralampos Pozidis
ISCAS2
2017 Fatiguing STDP: Learning from spike-timing codes in the presence of rate codes
abstract
Spiking neural networks (SNNs) could play a key role in unsupervised machine learning applications, by virtue of strengths related to learning from the fine temporal structure of event-based signals. However, some spike-timing-related strengths of SNNs are hindered by the sensitivity of spike-timing-dependent plasticity (STDP) rules to input spike rates, as fine temporal correlations may be obstructed by coarser correlations between firing rates. In this article, we propose a spike-timing-dependent learning rule that allows a neuron to learn from the temporally-coded information despite the presence of rate codes. Our long-term plasticity rule makes use of short-term synaptic fatigue dynamics. We show analytically that, in contrast to conventional STDP rules, our fatiguing STDP (FSTDP) helps learn the temporal code, and we derive the necessary conditions to optimize the learning process. We showcase the effectiveness of FSTDP in learning spike-timing correlations among processes of different rates in synthetic data. Finally, we use FSTDP to detect correlations in real-world weather data from the United States in an experimental realization of the algorithm that uses a neuro-morphic hardware platform comprising phase-change memristive devices. Taken together, our analyses and demonstrations suggest that FSTDP paves the way for the exploitation of the spike-based strengths of SNNs in real-world applications.
Timoleon Moraitis, Abu Sebastian, Irem Boybat, Manuel Le Gallo, Tomas Tuma, Evangelos Eleftheriou
IJCNN2
2011 Impulsive control for nanopositioning: stability and performance
abstract
In this paper, impulsive control is applied to a class of linear feedback systems and studied both theoretically and experimentally, with a particular focus on the usage in nanopositioning. By using impulsive control, improvements in tracking performance and tolerance to measurement noise can be achieved which are beyond the limits of conventional linear feedback.
Tomas Tuma, Angeliki Pantazi, John Lygeros, Abu Sebastian
HSCC4
2011 Programming algorithms for multilevel phase-change memory
abstract
Phase-change memory (PCM) has emerged as one among the most promising technologies for next-generation non-volatile solid-state memory. Multilevel storage, namely storage of non-binary information in a memory cell, is a key factor for reducing the total cost-per-bit and thus increasing the competiveness of PCM technology in the nonvolatile memory market. In this paper, we present a family of advanced programming schemes for multilevel storage in PCM. The proposed schemes are based on iterative write-and-verify algorithms that exploit the unique programming characteristics of PCM in order to achieve significant improvements in resistance-level packing density, robustness to cell variability, programming latency, energy- per-bit and cell storage capacity. Experimental results from PCM test-arrays are presented to validate the proposed programming schemes. In addition, the reliability issues of multilevel PCM in terms of resistance drift and read noise are discussed.
Nikolaos Papandreou, Haralampos Pozidis, Angeliki Pantazi, Abu Sebastian, Matthew J. Breitwisch, Chung Hon Lam, Evangelos Eleftheriou
ISCAS4
2010 Channel Modeling and Signal Processing for Probe Storage Channels
abstract
Probe-storage devices employ large arrays of probes to write/read data in parallel in some storage medium, and combine ultra-high density, low access times, and low power consumption. A particular probe-storage technique utilizes thermomechanical means to store and retrieve information in thin polymer films. In this paper, a system-level channel model for the thermomechanical probe-storage channel is presented. Each of the components of the proposed model is derived by extensive characterization of experimentally obtained readback signals from probe recording tests. Moreover, detection techniques that are actually utilized in a probe-storage prototype implementation are described, followed by coding techniques for added reliability in the presence of particles or other impurities of the storage medium. In addition to low-complexity coding constructs, a concatenated coding scheme with an outer LDPC and inner modulation code is considered, in order to establish a benchmark for overall system performance. A novel methodology for joint decoding of outer LDPC and inner (d,k) modulation codes is developed. Furthermore, an optimal soft decoder for the modulation code is proposed, based on a modification of the decoder metrics to accurately account for the probe storage channel output statistics. Experimental results are used throughout the paper to validate the channel model and identify its relevant parameters, as well as to verify the system performance obtained by simulations.
Haralampos Pozidis, Giovanni Cherubini, Angeliki Pantazi, Abu Sebastian, Evangelos Eleftheriou
IEEE J. Sel. Areas Commun.4
2007 Jitter Investigation and Performance Evaluation of a Small-Scale Probe Storage Device Prototype
abstract
MEMS-based scanning-probe data storage devices are emerging as potential ultra-high-density, low-access-time, and low-power alternatives to conventional data storage. Thermomechanical probe-based storage on thin polymer films is arguably the most advanced scanning-probe data storage scheme. The performance evaluation of a small-scale storage device prototype based on this concept is presented. The emphasis is on understanding the timing jitter in the read-back signals. Experiments are performed that confirm that the primary source of timing-jitter is the nanometer-scale perturbations of the micro-scanner while positioning the recording medium relative to the read/write transducers. Analytical estimates of these micro-scanner perturbations are obtained. An extensive performance evaluation, using the experimentally identified channel and medium-noise spectral characteristics, is conducted to study the impact of the microscanner perturbations on the performance of the storage device.
Abu Sebastian, Angeliki Pantazi, Haralampos Pozidis
GLOBECOM1
2005 Signal processing for probe storage
abstract
Scanning-probe data storage is emerging as a viable alternative to conventional data storage, offering ultra-high density, low access times, and low power consumption. One probe-storage technique utilizes a thermomechanical means to store and retrieve information in thin polymer films. We describe the readback signal path and characterize the thermomechanical-based probe-storage recording channel. It is shown that this channel exhibits a particular nonlinear behavior at high storage densities or high recording power, that is, the energy per unit time used to write a bit of information. A simple model is proposed that accurately captures the characteristics of this nonlinearity. Experimental results from single-probe recording setups are used to verify the validity of this model and identify its relevant parameters.
Haralampos Pozidis, Peter Bächtold, Giovanni Cherubini, Evangelos Eleftheriou, Christoph Hagleitner, Angeliki Pantazi, Abu Sebastian
ICASSP (5)7