Corey Lammie

dblp:228/3393 · also Corey Liam Lammie · DBLP profile ↗
← Back
19ranked-venue papers
11as first author
12since 2021 · last 2025
0000-0001-5564-1356ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 15 · 10 first-author · 8 since 2021Artificial intelligence and machine learning · 2 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2025 Hardware-Aware Compilation and Simulation for In-Memory Computing
abstract
This brief presents an overview of recent tools and research efforts aimed at enhancing the programmability and reliability of In-Memory Computing (IMC)-based systems. We discuss hardware-aware training techniques that improve model resilience to analog device imperfections, and explore mapping strategies that balance accuracy and performance for heterogeneous IMC-based accelerators. Additionally, we examine a compiler framework that abstracts hardware complexities and enables seamless integration of these accelerators into existing deployment pipelines. By combining these approaches with advanced simulation tools, we propose an end-to-end workflow that facilitates the practical deployment and optimization of IMC technologies across diverse memory types and architectural designs.
Asif Ali Khan, Hadjer Benmeziane, Hamid Farzaneh, João Paulo C. de Lima, William Andrew Simon, Yiyu Shi 0001, Zheyu Yan, Abu Sebastian, Xiaobo Sharon Hu, Jerónimo Castrillón, Corey Lammie
CASES11
2025 Assessing the Performance of Analog Training for Transfer Learning
abstract
Analog in-memory computing is a next-generation computing paradigm that promises fast, parallel, and energy-efficient deep learning training and transfer learning (TL). However, achieving this promise has remained elusive due to a lack of suitable training algorithms. Analog memory devices exhibit asymmetric and non-linear switching behavior in addition to device-to-device variation, meaning that most, if not all, of the current off-the-shelf training algorithms cannot achieve good training outcomes. Also, recently introduced algorithms have enjoyed limited attention, as they require bi-directionally switching devices of unrealistically high symmetry and precision and are highly sensitive. A new algorithm chopped TTv2 (c-TTv2), has been introduced, which leverages the chopped technique to address many of the challenges mentioned above. In this paper, we assess the performance of the c-TTv2 algorithm for analog TL using a Swin-ViT model on a subset of the CIFAR100 dataset. We also investigate the robustness of our algorithm to changes in some device specifications, including weight transfer noise, symmetry point skew, and symmetry point variability.
Omobayode Fagbohungbe, Corey Lammie, Malte J. Rasch, Takashi Ando, Tayfun Gokmen, Vijay Narayanan
ISCAS2
2025 Live Demonstration: Automated DNN Deployment on the IBM HERMES Project Chip
abstract
For this demonstration, we will showcase the operation of a software stack capable of automatically deploying Matrix-Vector Matrix (MVM) operations of diverse deep learning workloads in a pipelined-manner on a phase-change memory-based analog in-memory computing chip with high-accuracy. For a real chip, each deployment step will be highlighted for a transformer-based network trained to perform an organic chemical reaction prediction task. Additionally, using an emulated mode of operation, these steps will also be highlighted for a Resnet-based network, which has been trained to perform image classification, and a hybrid CNN/LSTM network trained to infer nucleotide sequences from sequences of amplitude values measured from a sequencing device.
Corey Lammie, Julian Büchel, Athanasios Vasilopoulos, Giacomo Camposampiero, Lionel Noussi, William Andrew Simon, Manuel Le Gallo, Abu Sebastian
ISCAS1
2025 A framework for analog-digital mixed-precision neural network training and inference
abstract
Recent advancements in AI hardware highlight the potential of mixed-signal accelerators, which integrate analog computation for matrix multiplications with reduced-precision digital operations, to achieve superior performance and energy efficiency. In this paper, we present a framework designed to perform hardware-aware training and inference evaluation of neural networks (NNs) on such accelerators. This framework extends an existing toolkit, the IBM Analog Hardware Acceleration Kit (AIHWKit), using a quantization library, enabling flexible layer-wise deployment in either analog or digital units, the latter with configurable precision and quantization options. Our combined framework supports simultaneous quantization-and analog-aware training as well as post-training calibration routines. It can also evaluate the accuracy of NNs when deployed on mixed-signal accelerators. We demonstrate the need of such a framework through ablation studies on a ResNet-based vision model and a BERT-based language model, highlighting the importance of its functionality for maximizing accuracy during deployment. Our contribution is open-sourced as part of the core code of AIHWKit [1].
Athanasios Vasilopoulos, Emma Boulharts, Corey Lammie, Julian Büchel, Hadjer Benmeziane, Manuel Le Gallo, Abu Sebastian
ISCAS3
2024 Analog AI as a Service: A Cloud Platform for In-Memory Computing
abstract
This paper introduces the Analog AI Cloud Composer platform, a service that allows users to access Analog In-Memory Computing (AIMC) simulation and computing resources over the cloud. We introduce the concept of an Analog AI as a Service (AAaaS). AIMC offers a novel approach for decreasing both the latency and energy usage associated with Deep Neural Network (DNN) inference and training. This platform democratizes access to AIMC computing, making it available to a broader audience, including researchers, developers, and businesses. Emphasizing a user-friendly, no-code approach, AAaaS integrates the Analog Hardware Acceleration Kit (AIHWKit) simulation platform within a fully managed cloud environment. We discuss the architecture of the Analog AI Cloud Composer (AAICC), focusing on its key services such as inference, training, and AIMC hardware access. The platform's design, grounded in cloud services and guidelines, ensures a secure, data-centric user experience with robust control and validation mechanisms.
Kaoutar El Maghraoui, Kim Tran, Kurtis Ruby, Borja Godoy, Jordan Murray, Manuel Le Gallo-Bourdeau, Todd Deshane, Pablo Gonzalez, Diego Moreda, Hadjer Benmeziane, Corey Lammie, Julian Büchel, Malte J. Rasch, Abu Sebastian, Vijay Narayanan
SSE11
2024 A Precision-Optimized Fixed-Point Near-Memory Digital Processing Unit for Analog In-Memory Computing
abstract
Analog In-Memory Computing (AIMC) is an emerging technology for fast and energy-efficient Deep Learning (DL) inference. However, a certain amount of digital post-processing is required to deal with circuit mismatches and non-idealities associated with the memory devices. Efficient near-memory digital logic is critical to retain the high area/energy efficiency and low latency of AIMC. Existing systems adopt Floating Point 16 (FP16) arithmetic with limited parallelization capability and high latency. To overcome these limitations, we propose a Near-Memory digital Processing Unit (NMPU) based on fixed-point arithmetic. It achieves competitive accuracy and higher computing throughput than previous approaches while minimizing the area overhead. Moreover, the NMPU supports standard DL activation steps, such as ReLU and Batch Normalization. We perform a physical implementation of the NMPU design in a 14 nm CMOS technology and provide detailed performance, power, and area assessments. We validate the efficacy of the NMPU by using data from an AIMC chip and demonstrate that a simulated AIMC system with the proposed NMPU outperforms existing FP16- based implementations, providing 139 × speed-up, 7.8 × smaller area, and a competitive power consumption. Additionally, our approach achieves an inference accuracy of 86.65 %/65.06 %, with an accuracy drop of just 0.12 %/0.4 % compared to the FP16 baseline when benchmarked with ResNet9/ResNet32 networks trained on the CIFAR10/CIFAR100 datasets, respectively.
Elena Ferro, Athanasios Vasilopoulos, Corey Lammie, Manuel Le Gallo, Luca Benini, Irem Boybat, Abu Sebastian
ISCAS3
2024 Improving the Accuracy of Analog-Based In-Memory Computing Accelerators Post-Training
abstract
Analog-Based In-Memory Computing (AIMC) inference accelerators can be used to efficiently execute Deep Neural Network (DNN) inference workloads. However, to mitigate accuracy losses, due to circuit and device non-idealities, Hardware-Aware (HWA) training methodologies must be employed. These typically require significant information about the underlying hardware. In this paper, we propose two Post-Training (PT) optimization methods to improve accuracy after training is performed. For each crossbar, the first optimizes the conductance range of each column, and the second optimizes the input, i.e, Digital-to-Analog Converter (DAC), range. It is demonstrated that, when these methods are employed, the complexity during training, and the amount of information about the underlying hardware can be reduced, with no notable change in accuracy (≤0.1%) when finetuning the pretrained RoBERTa transformer model for all General Language Understanding Evaluation (GLUE) benchmark tasks. Additionally, it is demonstrated that further optimizing learned parameters PT improves accuracy.
Corey Lammie, Athanasios Vasilopoulos, Julian Büchel, Giacomo Camposampiero, Manuel Le Gallo, Malte J. Rasch, Abu Sebastian
ISCAS1
2024 Unsupervised character recognition with graphene memristive synapses
Ben Walters, Corey Lammie, Shuangming Yang, Mohan V. Jacob, Mostafa Rahimi Azghadi
Neural Comput. Appl.2
2022 Design Space Exploration of Dense and Sparse Mapping Schemes for RRAM Architectures
abstract
The impact of device and circuit-level effects in mixed-signal Resistive Random Access Memory (RRAM) accelerators typically manifest as performance degradation of Deep Learning (DL) algorithms, but the degree of impact varies based on algorithmic features. These include network architecture, capacity, weight distribution, and the type of inter-layer connections. Techniques are continuously emerging to efficiently train sparse neural networks, which may have activation sparsity, quantization, and memristive noise. In this paper, we present an extended Design Space Exploration (DSE) methodology to quantify the benefits and limitations of dense and sparse mapping schemes for a variety of network architectures. While sparsity of connectivity promotes less power consumption and is often optimized for extracting localized features, its performance on tiled RRAM arrays may be more susceptible to noise due to under-parameterization, when compared to dense mapping schemes. Moreover, we present a case study quantifying and formalizing the trade-offs of typical non-idealities introduced into l-Transistor-l-Resistor (ITIR) tiled memristive architectures and the size of modular crossbar tiles using the CIFAR-10 dataset.
Corey Lammie, Jason Kamran Eshraghian, Chenqi Li, Amirali Amirsoleimani, Roman Genov, Wei Lu 0003, Mostafa Rahimi Azghadi
ISCAS1
2022 MemTorch: An Open-source Simulation Framework for Memristive Deep Learning Systems
Corey Lammie, Wei Xiang 0001, Bernabé Linares-Barranco, Mostafa Rahimi Azghadi
Neurocomputing1
2021 Towards Memristive Deep Learning Systems for Real-Time Mobile Epileptic Seizure Prediction
abstract
The unpredictability of seizures continues to distress many people with drug-resistant epilepsy. On account of recent technological advances, considerable efforts have been made using different hardware technologies to realize smart devices for the real-time detection and prediction of seizures. In this paper, we investigate the feasibility of using Memristive Deep Learning Systems (MDLSs) to perform real-time epileptic seizure prediction on the edge. Using the MemTorch simulation framework and the Children's Hospital Boston (CHB)-Massachusetts Institute of Technology (MIT) dataset we determine the performance of various simulated MDLS configurations. An average sensitivity of 77.4% and a Area Under the Receiver Operating Characteristic Curve (AUROC) of 0.85 are reported for the optimal configuration that can process Electroencephalogram (EEG) spectrograms with 7,680 samples in 1.408ms while consuming 0.0133W and occupying an area of 0.1269mm2in a 65nm Complementary Metal-Oxide-Semiconductor (CMOS) process.
Corey Lammie, Wei Xiang 0001, Mostafa Rahimi Azghadi
ISCAS1
2021 A Deep Learning Localization Method for Measuring Abdominal Muscle Dimensions in Ultrasound Images
abstract
Health professionals extensively use Two-Dimensional (2D) Ultrasound (US) videos and images to visualize and measure internal organs for various purposes including evaluation of muscle architectural changes. US images can be used to measure abdominal muscles dimensions for the diagnosis and creation of customized treatment plans for patients with Low Back Pain (LBP), however, they are difficult to interpret. Due to high variability, skilled professionals with specialized training are required to take measurements to avoid low intra-observer reliability. This variability stems from the challenging nature of accurately finding the correct spatial location of measurement endpoints in abdominal US images. In this paper, we use a Deep Learning (DL) approach to automate the measurement of the abdominal muscle thickness in 2D US images. By treating the problem as a localization task, we develop a modified Fully Convolutional Network (FCN) architecture to generate blobs of coordinate locations of measurement endpoints, similar to what a human operator does. We demonstrate that using the TrA400 US image dataset, our network achieves a Mean Absolute Error (MAE) of 0.3125 on the test set, which almost matches the performance of skilled ultrasound technicians. Our approach can facilitate next steps for automating the process of measurements in 2D US images, while reducing inter-observer as well as intra-observer variability for more effective clinical outcomes.
Alzayat Saleh, Issam H. Laradji, Corey Lammie, David Vázquez 0001, Carol A. Flavell, Mostafa Rahimi Azghadi
IEEE J. Biomed. Health Informatics3
2020 Biologically Plausible Contrast Detection using a Memristor Array
abstract
Hardware implementation of functional neuronal circuits has rapidly become more feasible due to increasing reliability of memristor-CMOS integration at a scale necessitated by neuromorphic processes. Most neuromorphic implementations of the memristor treat it as a variable synaptic weight modulated by conductance. The work in this paper enhances biological plausibility of analog vision system circuits by mimicking the nonlinear dynamics of a network of receptive field that resembles those found in the lateral geniculate nucleus. The memristive circuit provides a biologically accurate response where a single cell behaves simultaneously as its own on-center receptive field, and as the off-surround receptive field of adjacent cells. It maximally responds to spatial variations of light. Each output is a fundamental unit of cortical visual information, and we show how the receptive field of each neuron can be superimposed to perceive edges and recognize objects when scaled to higher cortical areas. The functionality of the array is verified in SPICE simulations.
Jason Kamran Eshraghian, Corey Lammie, Mostafa Rahimi Azghadi
ISCAS2
2020 Live Demonstration: Low-Power and High-Speed Deep FPGA Inference Engines for Weed Classification at the Edge
abstract
In Low-Power and High-Speed Deep FPGA Inference Engines for Weed Classification at the Edge [1] we implemented GPU- and FPGA-accelerated deterministically binarized Deep Neural Networks (DNNs), tailored toward weed species classification for robotic weed control. The dataset used consisted of 17,508 unique 256×256 color images in 9 classes, collected in situ from eight rangeland areas across Northern Australia [2]. For this live demonstration, we have designed a weed classification game. We first provide the visitor with a printed sheet showing several examples of each of the 9 various weed species classes in our dataset, to learn and memorize the weed names. This learning process can take for as long as the visitor wishes. For the game to start, five weed images from our test set are randomly selected. We then measure the interference times and accuracies for our optimized GPUand FPGA-accelerated binarized DNNs, alongside the visitors' performance. Are low-resolution, low-power, binarized DNNs able to outperform humans at categorizing weed species?
Corey Lammie, Mostafa Rahimi Azghadi
ISCAS1
2020 MemTorch: A Simulation Framework for Deep Memristive Cross-Bar Architectures
abstract
Memristive devices arranged in cross-bar architectures have shown great promise to facilitate the acceleration and improve the power efficiency of Deep Learning (DL) systems for deployment in resource-constrained platforms, such as the Internet-of-Things (IoT) edge devices. These cross-bar architectures can be used to implement various in-memory computing operations, such as Multiply-Accumulate (MAC) and convolution, which are used extensively in Deep Neural Networks (DNNs) and Convolutional Neural Networks (CNNs). Currently, there is a lack of an open source, general, high-level simulation platform that can fully integrate any behavioral or experimental memristive device model into cross-bar architectures. This paper presents such a framework named MemTorch, which integrates directly with the well-known PyTorch Machine Learning (ML) library. To demonstrate an example practical use of MemTorch, we use it to simulate the performance degradation that non-ideal devices introduce to a typical Memristive DNN (MDNN) implementing VGG-16 for CIFAR-10. Our open source1MemTorch framework can be used by circuit and system designers to conveniently build customized large-scale simulation platforms, as a preliminary step before circuit-level realization.
Corey Lammie, Mostafa Rahimi Azghadi
ISCAS1
2020 Training Progressively Binarizing Deep Networks using FPGAs
abstract
While hardware implementations of inference routines for Binarized Neural Networks (BNNs) are plentiful, current realizations of efficient BNN hardware training accelerators, suitable for Internet of Things (IoT) edge devices, leave much to be desired. Conventional BNN hardware training accelerators perform forward and backward propagations with parameters adopting binary representations, and optimization using parameters adopting floating or fixed-point real-valued representations-requiring two distinct sets of network parameters. In this paper, we propose a hardware-friendly training method that, contrary to conventional methods, progressively binarizes a singular set of fixed-point network parameters, yielding notable reductions in power and resource utilizations. We use the Intel FPGA SDK for OpenCL development environment to train our progressively binarizing DNNs on an OpenVINO FPGA. We benchmark our training approach on both GPUs and FPGAs using CIFAR-10 and compare it to conventional BNNs.
Corey Lammie, Wei Xiang 0001, Mostafa Rahimi Azghadi
ISCAS1
2019 Stochastic Computing for Low-Power and High-Speed Deep Learning on FPGA
abstract
Stochastic Computing (SC) presents a low-cost and low-power alternative to conventional binary computing. In SC, continuous values are represented by stochastically generated bit streams. By performing simple hardware-friendly bit-wise operations on these streams, complex calculations can be realized very efficiently. However, the inherent randomness and approximation used in SC can result in undesirable computational errors. As Convolutional Neural Networks (CNNs) are inherently error-tolerant, SC could be embedded in them to gain higher speed and lower power without significant accuracy loss. In this paper, we propose using SC techniques to approximate multiplication operations on fixed-point weights and biases during training of CNNs. By employing such techniques, we demonstrate near state-of-the-art learning performance for the MNIST and CIFAR-10 datasets, while achieving significant resource and speed improvements when implementing the deep networks on a Field Programmable Gate Array (FPGA). For MNIST, we demonstrate that SC compared to conventional computing, will result in almost 3 times increase in learning speed with only 1.37% degradation in validation accuracy. Similarly, for CIFAR-10, training is accelerated 3.5 times with a degradation of 3.39%. We also show that our FPGA implementations of CNNs adopting stochastic multipliers consume over 17 times less power than their GPU counterparts.
Corey Lammie, Mostafa Rahimi Azghadi
ISCAS1
2018 Live Demonstration: Unsupervised Character Recognition with a FPGA Neuromorphic System
abstract
For this demonstration, we have implemented a Spiking Neural Network (SNN) on a Field Programmable Gate Array (FPGA) and trained it using Spike Timing Dependent Plasticity (STDP) to identify temporally encoded characters, in an unsupervised manner. The constructed one-layer network consists of plastic excitatory and non-plastic inhibitory synapses, which are connected to output Izhikevich neurons. The implemented neural hardware demonstrates a powerful and fast learning scheme, which brings about a significant unsupervised classification accuracy of 94 %.
Corey Lammie, Tara J. Hamilton, Mostafa Rahimi Azghadi
ISCAS1
2018 Unsupervised Character Recognition with a Simplified FPGA Neuromorphic System
abstract
Neuromorphic hardware platforms have demonstrated significant promise in cognitive tasks such as visual processing and classification. These platforms usually consist of several layers of spiking neurons for feature extraction and various learning mechanisms, which renders the associated networks power and hardware hungry. In this paper, we have implemented a simplified proof-of-concept Spiking Neural Network (SNN) on a Field Programmable Gate Array (FPGA) and trained it using Spike Timing Dependent Plasticity (STDP) to identify temporally encoded characters, in an unsupervised manner. The constructed one-layer network consists of excitatory synapses, which receive input characters in the form of Poissonian spike trains from the pre-synaptic side. From the post-synaptic side, the synapses are connected to output Izhikevich neurons. In addition, non-plastic inhibitory synapses between the output neurons are introduced to implement lateral inhibition and competitive learning. The implemented neural hardware demonstrates a powerful and fast learning scheme, which brings about a significant unsupervised classification accuracy of 94 %. In addition, since the proposed network receives the characters in the form of spike trains, it is amenable to being interfaced to neuromorphic event-driven sensors such as silicon retina, making the proposed platform useful for online unsupervised template matching applications.
Corey Lammie, Tara J. Hamilton, Mostafa Rahimi Azghadi
ISCAS1