Charles Mackin

dblp:130/1390 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
3since 2021 · last 2023
0000-0001-8413-5583ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 3 since 2021
YearPublicationVenuePosition
2023 Architectures and Circuits for Analog-memory-based Hardware Accelerators for Deep Neural Networks (Invited)
abstract
Analog non-volatile memory (NVM)-based accelerators for Deep Neural Networks (DNNs) can achieve high-throughput and energy-efficient multiply-accumulate (MAC) operations by taking advantage of massively parallelized analog compute, implemented with Ohm's law and Kirchhoff's current law on arrays of resistive memory devices. Competitive end-to-end DNN accuracies can be obtained, provided that weights are accurately programmed onto NVM devices and MAC operations are sufficiently linear. In this paper, we report architectural and circuit advances for such Analog NVM-based accelerators. We describe a highly heterogeneous and programmable accelerator architecture for DNN inference that combines analog NVM memory-array “Tiles” for weight-stationary, energy-efficient MAC operations, together with heterogeneous special-function Compute-Cores for auxiliary digital computation. Massively parallel vectors of neuron-activation data are exchanged over short distances using a dense and efficient circuit-switched 2D mesh, enabling a wide range of DNN workloads, including CNNs, LSTMs, and Transformers. We also show a 14-nm inference chip consisting of multiple$\mathbf{512}\times \mathbf{512}$arrays of Phase Change Memory (PCM) devices which implements multiple DNN benchmarks using such a circuit-switched 2D mesh.
Hsinyu Tsai, Pritish Narayanan, Shubham Jain 0004, Stefano Ambrogio, Kohji Hosokawa, Masatoshi Ishii, Charles Mackin, Ching-Tzu Chen, Atsuya Okazaki, Akiyo Nomura, Irem Boybat, Ramachandran Muralidhar, Martin M. Frank, Takeo Yasuda, Alexander M. Friz, Yasuteru Kohda, An Chen 0002, Andrea Fasoli, Malte J. Rasch, Stanislaw Wozniak, Jose Luquin, Vijay Narayanan, Geoffrey W. Burr
ISCAS7
2022 Analog-memory-based 14nm Hardware Accelerator for Dense Deep Neural Networks including Transformers
abstract
Analog non-volatile memory (NVM)-based accelerators for deep neural networks perform high-throughput and energy-efficient multiply-accumulate (MAC) operations (e.g., high TeraOPS/W) by taking advantage of massively parallelized analog MAC operations, implemented with Ohm’s law and Kirchhoff’s current law on array-matrices of resistive devices. While the wide-integer and floating-point operations offered by conventional digital CMOS computing are much more suitable than analog computing for conventional applications that require high accuracy and true reproducibility, deep neural networks can still provide competitive end-to-end results even with modest (e.g., 4-bit) precision in synaptic operations. In this paper, we describe a 14-nm inference chip, comprising multiple 512$\times$ 512 arrays of Phase Change Memory (PCM) devices, which can deliver software-equivalent inference accuracy for MNIST handwritten-digit recognition and recurrent LSTM benchmarks, by using compensation techniques to finesse analog-memory challenges such as conductance drift and noise. We also project accuracy for Natural Language Processing (NLP) tasks performed with a state-of-art large Transformer-based model, BERT, when mapped onto an extended version of this same fundamental chip architecture.
Atsuya Okazaki, Pritish Narayanan, Stefano Ambrogio, Kohji Hosokawa, Hsinyu Tsai, Akiyo Nomura, Takeo Yasuda, Charles Mackin, Alexander M. Friz, Masatoshi Ishii, Yasuteru Kohda, Katie Spoon, An Chen 0002, Andrea Fasoli, Malte J. Rasch, Geoffrey W. Burr
ISCAS8
2021 Circuit Techniques for Efficient Acceleration of Deep Neural Network Inference with Analog-AI (Invited)
abstract
By performing parallelized multiply-accumulate operations in the analog domain at the location of weight data, crossbar-array "tiles" of analog non-volatile memory (NVM) devices can potentially accelerate the forward-inference of deep neural networks. To be successful, such systems will need to achieve two related but challenging goals. First is the achievement of high neural network classification accuracies, indistinguishable from those achieved with conventional approaches, despite the difficulties of programming NVM devices accurately in the presence of significant device-to-device variability. Towards this first goal, we describe row-wise Phase-Change Memory (PCM) programming schemes for rapid yet accurate weight- programming. The second goal is highly energy-efficient forward- inference of multi-layer neural networks, requiring efficiency in both the massively-parallel analog-AI operations performed at each tile, as well as efficiency in how the resulting neuron-excitation data vectors get conveyed from tile to tile. Towards this second goal, micro-architectural design ideas including source-follower-based readout, array segmentation, and transmit-by- duration are described.
Kohji Hosokawa, Pritish Narayanan, Stefano Ambrogio, Hsinyu Tsai, Charles Mackin, Andrea Fasoli, Alexander M. Friz, An Chen 0002, Jose Luquin, Katie Spoon, Geoffrey W. Burr, Scott C. Lewis
ISCAS5
2020 Optimization of Analog Accelerators for Deep Neural Networks Inference
abstract
Neuromorphic computation based on analog nonvolatile memories (NVMs) holds great promise to improve Deep Neural Networks inference performance. In virtue of an architecture that executes the computation at the location of the stored weight data, remarkable gains in energy efficiency and speed are projected over competing von Neumann architectures leveraged by existing digital accelerators. Here we describe two optimization strategies for NVMs: one for programming the memory elements and one of cell design, both aimed at mitigating the effect of NVM non-idealities on the performance of analog, phase-change memory-based accelerators. We then demonstrate the advantages realized by such strategies on the inference accuracy of Long Short Term Memory networks and evaluate the energy requirements of such networks.
Andrea Fasoli, Stefano Ambrogio, Pritish Narayanan, Hsinyu Tsai, Charles Mackin, Katie Spoon, Alexander M. Friz, An Chen 0002, Geoffrey W. Burr
ISCAS5
2013 Rapid exploration of processing and design guidelines to overcome carbon nanotube variations
abstract
Carbon nanotube field-effect transistors (CNFETs) are promising candidates for building energy-efficient digital systems at highly-scaled technology nodes. However, carbon nanotubes (CNTs) are inherently subject to variations that reduce circuit yield, increase susceptibility to noise, and severely degrade their anticipated energy and speed benefits. Joint exploration and optimization of CNT processing options and CNFET circuit design are required to overcome this outstanding challenge. Unfortunately, existing approaches for such exploration and optimization are computationally expensive, and mostly rely on trial-and-error-based ad-hoc techniques. In this paper, we present a systematic framework which quickly evaluates the impact of CNT variations on circuit delay and noise margin, and automatically explores the large space of CNT processing options to derive optimized CNT processing and CNFET circuit design guidelines. We demonstrate that: 1. Our new framework runs over 100X faster than existing approaches. 2. It accurately identifies the most important CNT processing parameters, together with CNFET circuit sizing, to minimize the impact of CNT variations while meeting circuit-level noise margin constraints.
Gage Hills, Jie Zhang 0007, Charles Mackin, Max M. Shulaker, Hai Wei, H.-S. Philip Wong, Subhasish Mitra
DAC3