Geoffrey W. Burr

dblp:67/3435 · DBLP profile ↗
← Back
11ranked-venue papers
1as first author
7since 2021 · last 2025
0000-0001-5717-2549ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 11 · 1 first-author · 7 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
YearPublicationVenuePosition
2025 NORA: Noise-Optimized Rescaling of LLMs on Analog Compute-in-Memory Accelerators
abstract
Large Language Models (LLMs) have become critical in AI applications, yet current digital AI accelerators suffer from significant energy inefficiencies due to frequent data movement. Analog compute-in-memory (CIM) accelerators offer a potential solution for improving energy efficiency but introduce non-idealities that can degrade LLM accuracy. While analog CIM has been extensively studied for traditional deep neural networks, its impact on LLMs remains unexplored, particularly concerning the large influence of Analog CIM non-idealities. In this paper, we conduct a sensitivity analysis on the effects of analog-induced noise on LLM accuracy. We find that while LLMs demonstrate robustness to weight-related noise, they are highly sensitive to quantization noise and additive Gaussian noise. Based on these insights, we propose a noise-optimized rescaling method to mitigate LLM accuracy loss by shifting the non-ideality burden from the sensitive input/output to the more resilient weight. Through rescaling, we can implement the OPT-6.7b model on simulated analog CIM hardware with less than 1% accuracy loss from the floating-point baseline, compared to a much higher loss of around 30% without rescaling.
Yayue Hou, Hsinyu Tsai, Kaoutar El Maghraoui, Tayfun Gokmen, Geoffrey W. Burr, Liu Liu 0017
DATE5
2025 SAGE: Saliency-Aware Grouping for Efficient Mapping of LLMs on Analog Compute-in-Memory
abstract
Large Language Models (LLMs) demand high memory bandwidth and computational efficiency, posing significant challenges for deployment on traditional digital accelerators. Analog Compute-in-Memory (ACIM) architectures offer an attractive alternative by co-locating storage and computation to reduce data movement. However, executing LLMs on ACIM systems remains challenging due to hardware non-idealities and the unique statistical properties of LLM inputs and outputs in FC layers. In particular, long-tailed data distributions containing large-amplitude "salient values" degrade analog signal quality under quantization and system noise. In this work, we propose SAGE (Saliency-Aware Grouping for Efficient Mapping), a training-free strategy that improves noise resilience by reordering weight and input channels of FC layers based on statistical characteristics of LLMs. We identify kurtosis as a key factor affecting analog robustness and develop a saliency-aware mapping method that reduces output kurtosis to enhance the signal-to-noise ratio. We further introduce a reconfigurable tile design that supports mixed-precision execution and maximizes array utilization across layers. Evaluations on multiple LLMs and benchmarks show that SAGE significantly improves inference accuracy and energy efficiency based on ACIM simulation without requiring retraining.
Yayue Hou, Garrett Gagnon, Hsinyu Tsai, Kaoutar El Maghraoui, Geoffrey W. Burr, Liu Liu 0017
ICCAD6
2025 CiMBA: Accelerating Genome Sequencing Through On-Device Basecalling via Compute-in-Memory
abstract
As genome sequencing is finding utility in a wide variety of domains beyond the confines of traditional medical settings, its computational pipeline faces two significant challenges. First, the creation of up to 0.5 GB of data per minute imposes substantial communication and storage overheads. Second, the sequencing pipeline is bottlenecked at the basecalling step, consuming >40% of genome analysis time. A range of proposals have attempted to address these challenges, with limited success. We propose to address these challenges with a Compute-in-Memory Basecalling Accelerator (CiMBA), the first embedded ($\sim 25$mm$^{2}$) accelerator capable of real-time, on-device basecalling, coupled with AnaLog (AL)-Dorado, a new family of analog focused basecalling DNNs. Our resulting hardware/software co-design greatly reduces data communication overhead, is capable of a throughput of 4.77 million bases per second, 24× that required for real-time operation, and achieves 17 × /27× power/area efficiency over the best prior basecalling embedded accelerator while maintaining a high accuracy comparable to state-of-the-art software basecallers.
William Andrew Simon, Irem Boybat, Riselda Kodra, Elena Ferro, Gagandeep Singh 0002, Mohammed Alser, Shubham Jain 0004, Hsinyu Tsai, Geoffrey W. Burr, Onur Mutlu, Abu Sebastian
IEEE Trans. Parallel Distributed Syst.9
2023 Architectures and Circuits for Analog-memory-based Hardware Accelerators for Deep Neural Networks (Invited)
abstract
Analog non-volatile memory (NVM)-based accelerators for Deep Neural Networks (DNNs) can achieve high-throughput and energy-efficient multiply-accumulate (MAC) operations by taking advantage of massively parallelized analog compute, implemented with Ohm's law and Kirchhoff's current law on arrays of resistive memory devices. Competitive end-to-end DNN accuracies can be obtained, provided that weights are accurately programmed onto NVM devices and MAC operations are sufficiently linear. In this paper, we report architectural and circuit advances for such Analog NVM-based accelerators. We describe a highly heterogeneous and programmable accelerator architecture for DNN inference that combines analog NVM memory-array “Tiles” for weight-stationary, energy-efficient MAC operations, together with heterogeneous special-function Compute-Cores for auxiliary digital computation. Massively parallel vectors of neuron-activation data are exchanged over short distances using a dense and efficient circuit-switched 2D mesh, enabling a wide range of DNN workloads, including CNNs, LSTMs, and Transformers. We also show a 14-nm inference chip consisting of multiple$\mathbf{512}\times \mathbf{512}$arrays of Phase Change Memory (PCM) devices which implements multiple DNN benchmarks using such a circuit-switched 2D mesh.
Hsinyu Tsai, Pritish Narayanan, Shubham Jain 0004, Stefano Ambrogio, Kohji Hosokawa, Masatoshi Ishii, Charles Mackin, Ching-Tzu Chen, Atsuya Okazaki, Akiyo Nomura, Irem Boybat, Ramachandran Muralidhar, Martin M. Frank, Takeo Yasuda, Alexander M. Friz, Yasuteru Kohda, An Chen 0002, Andrea Fasoli, Malte J. Rasch, Stanislaw Wozniak, Jose Luquin, Vijay Narayanan, Geoffrey W. Burr
ISCAS23
2023 A Heterogeneous and Programmable Compute-In-Memory Accelerator Architecture for Analog-AI Using Dense 2-D Mesh
abstract
We introduce a highly heterogeneous and programmable compute-in-memory (CIM) accelerator architecture for deep neural network (DNN) inference. This architecture combines spatially distributed CIM memory array “tiles” for weight-stationary, energy-efficient multiply–accumulate (MAC) operations, together with heterogeneous special-function compute cores for auxiliary digital computation. Massively parallel vectors of neuron activation data are exchanged over short distances using a dense and efficient circuit-switched 2-D mesh, offering full end-to-end support for a wide range of DNN workloads, including CNNs, long-short-term-memory (LSTM), and transformers. We discuss the design of the “analog fabric”—the 2-D grid of tiles and compute cores interconnected by the 2-D mesh—and address the efficiency in both mapping of DNNs onto the hardware and in pipelining of various DNN workloads across a range of batch sizes. We show, for the first time, system-level assessments using projected component parameters for a realistic “analog AI” system, based on dense crossbar arrays of low-power nonvolatile analog memory elements, while incorporating a single common analog fabric design that can scale to large networks by introducing data transport between multiple analog AI chips. Our performance estimates for several networks, including large LSTM and bidirectional encoder representations from transformers (BERT), show highly competitive throughput while offering$40\times $–$140\times $higher energy efficiency than NVIDIA A100—thus illustrating the strong promise of analog AI and the proposed architecture for DNN inference applications.
Shubham Jain 0004, Hsinyu Tsai, Ching-Tzu Chen, Ramachandran Muralidhar, Irem Boybat, Martin M. Frank, Stanislaw Wozniak, Milos Stanisavljevic, Praneet Adusumilli, Pritish Narayanan, Kohji Hosokawa, Masatoshi Ishii, Vijay Narayanan, Geoffrey W. Burr
IEEE Trans. Very Large Scale Integr. Syst.15
2022 Analog-memory-based 14nm Hardware Accelerator for Dense Deep Neural Networks including Transformers
abstract
Analog non-volatile memory (NVM)-based accelerators for deep neural networks perform high-throughput and energy-efficient multiply-accumulate (MAC) operations (e.g., high TeraOPS/W) by taking advantage of massively parallelized analog MAC operations, implemented with Ohm’s law and Kirchhoff’s current law on array-matrices of resistive devices. While the wide-integer and floating-point operations offered by conventional digital CMOS computing are much more suitable than analog computing for conventional applications that require high accuracy and true reproducibility, deep neural networks can still provide competitive end-to-end results even with modest (e.g., 4-bit) precision in synaptic operations. In this paper, we describe a 14-nm inference chip, comprising multiple 512$\times$ 512 arrays of Phase Change Memory (PCM) devices, which can deliver software-equivalent inference accuracy for MNIST handwritten-digit recognition and recurrent LSTM benchmarks, by using compensation techniques to finesse analog-memory challenges such as conductance drift and noise. We also project accuracy for Natural Language Processing (NLP) tasks performed with a state-of-art large Transformer-based model, BERT, when mapped onto an extended version of this same fundamental chip architecture.
Atsuya Okazaki, Pritish Narayanan, Stefano Ambrogio, Kohji Hosokawa, Hsinyu Tsai, Akiyo Nomura, Takeo Yasuda, Charles Mackin, Alexander M. Friz, Masatoshi Ishii, Yasuteru Kohda, Katie Spoon, An Chen 0002, Andrea Fasoli, Malte J. Rasch, Geoffrey W. Burr
ISCAS16
2021 Circuit Techniques for Efficient Acceleration of Deep Neural Network Inference with Analog-AI (Invited)
abstract
By performing parallelized multiply-accumulate operations in the analog domain at the location of weight data, crossbar-array "tiles" of analog non-volatile memory (NVM) devices can potentially accelerate the forward-inference of deep neural networks. To be successful, such systems will need to achieve two related but challenging goals. First is the achievement of high neural network classification accuracies, indistinguishable from those achieved with conventional approaches, despite the difficulties of programming NVM devices accurately in the presence of significant device-to-device variability. Towards this first goal, we describe row-wise Phase-Change Memory (PCM) programming schemes for rapid yet accurate weight- programming. The second goal is highly energy-efficient forward- inference of multi-layer neural networks, requiring efficiency in both the massively-parallel analog-AI operations performed at each tile, as well as efficiency in how the resulting neuron-excitation data vectors get conveyed from tile to tile. Towards this second goal, micro-architectural design ideas including source-follower-based readout, array segmentation, and transmit-by- duration are described.
Kohji Hosokawa, Pritish Narayanan, Stefano Ambrogio, Hsinyu Tsai, Charles Mackin, Andrea Fasoli, Alexander M. Friz, An Chen 0002, Jose Luquin, Katie Spoon, Geoffrey W. Burr, Scott C. Lewis
ISCAS11
2020 Optimization of Analog Accelerators for Deep Neural Networks Inference
abstract
Neuromorphic computation based on analog nonvolatile memories (NVMs) holds great promise to improve Deep Neural Networks inference performance. In virtue of an architecture that executes the computation at the location of the stored weight data, remarkable gains in energy efficiency and speed are projected over competing von Neumann architectures leveraged by existing digital accelerators. Here we describe two optimization strategies for NVMs: one for programming the memory elements and one of cell design, both aimed at mitigating the effect of NVM non-idealities on the performance of analog, phase-change memory-based accelerators. We then demonstrate the advantages realized by such strategies on the inference accuracy of Long Short Term Memory networks and evaluate the energy requirements of such networks.
Andrea Fasoli, Stefano Ambrogio, Pritish Narayanan, Hsinyu Tsai, Charles Mackin, Katie Spoon, Alexander M. Friz, An Chen 0002, Geoffrey W. Burr
ISCAS9
2017 Neuromorphic devices and architectures for next-generation cognitive computing
abstract
Cognitive computing describes “systems that learn at scale, reason with purpose, and interact with humans naturally” [1]. In this paper, we review our work towards enabling “next generation” cognitive computing using neuromorphic computational schemes that could potentially outperform present-day CPUs and GPUs. Here we use large arrays of Resistive Non-Volatile Memories (NVM) with device conductance serving as synaptic weight. We focus on training and classification using fully-connected networks based on the backpropagation algorithm, and show that our approach could offer power and speed advantages over conventional Von-Neumann processors. We also propose some circuit approximations that improve network parallelism without significantly degrading classification accuracy. Finally, we explore the requirements for a system implementation of on-chip learning.
Geoffrey W. Burr, Pritish Narayanan, Robert M. Shelby, Stefano Ambrogio, Hsinyu Tsai, Scott L. Lewis, Kohji Hosokawa
ISCAS1
2017 Reducing circuit design complexity for neuromorphic machine learning systems based on Non-Volatile Memory arrays
abstract
Machine Learning (ML) is an attractive application of Non-Volatile Memory (NVM) arrays [1,2]. However, achieving speedup over GPUs will require minimal neuron circuit sharing and thus highly area-efficient peripheral circuitry, so that ML reads and writes are massively parallel and time-multiplexing is minimized [2]. This means that neuron hardware offering full `software-equivalent' functionality is impractical. We analyze neuron circuit needs for implementing back-propagation in NVM arrays and introduce approximations to reduce design complexity and area. We discuss the interplay between circuits and NVM devices, such as the need for an occasional RESET step, the number of programming pulses to use, and the stochastic nature of NVM conductance change. In all cases we show that by leveraging the resilience of the algorithm to error, we can use practical circuit approaches yet maintain competitive test accuracies on ML benchmarks.
Pritish Narayanan, Lucas L. Sanches, Alessandro Fumarola, Robert M. Shelby, Stefano Ambrogio, Jun-Woo Jang, Hyunsang Hwang, Yusuf Leblebici, Geoffrey W. Burr
ISCAS9
2013 Nanoscale electronic synapses using phase change devices
abstract
The memory capacity, computational power, communication bandwidth, energy consumption, and physical size of the brain all tend to scale with the number of synapses, which outnumber neurons by a factor of 10,000. Although progress in cortical simulations using modern digital computers has been rapid, the essential disparity between the classical von Neumann computer architecture and the computational fabric of the nervous system makes large-scale simulations expensive, power hungry, and time consuming. Over the last three decades, CMOS-based neuromorphic implementations of “electronic cortex” have emerged as an energy efficient alternative for modeling neuronal behavior. However, the key ingredient for electronic implementation of any self-learning system—programmable, plastic Hebbian synapses scalable to biological densities—has remained elusive. We demonstrate the viability of implementing such electronic synapses using nanoscale phase change devices. We introduce novel programming schemes for modulation of device conductance to closely mimic the phenomenon of Spike Timing Dependent Plasticity (STDP) observed biologically, and verify through simulations that such plastic phase change devices should support simple correlative learning in networks of spiking neurons. Our devices, when arranged in a crossbar array architecture, could enable the development of synaptronic systems that approach the density (∼10 11 synapses per sq cm) and energy efficiency (consuming ∼1pJ per synaptic programming event) of the human brain.
Bryan L. Jackson, Bipin Rajendran, Gregory S. Corrado, Matthew J. Breitwisch, Geoffrey W. Burr, Roger Cheek, Kailash Gopalakrishnan, Simone Raoux, Charles T. Rettner, Alvaro Padilla, Alejandro G. Schrott, Rohit S. Shenoy, Bülent N. Kurdi, Chung Hon Lam, Dharmendra S. Modha
ACM J. Emerg. Technol. Comput. Syst.5